ae099482a9e18c1987df1f32534b2b1791c6d978
63
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ae099482a9 |
011: twenty cases, and two axes nothing covers
Before settling a schema, what a module can actually be. Twenty kinds of thing, with the hard ones at the end because they are the point. The ordinary nine are unsurprising: a supervised service, a system package with configuration, an application a person launches, a command-line tool, a library that never runs, a one-shot task, a scheduled one, an adapter, and a standalone application whose only difference is where its source lives. The eleven that break a naive schema are where the work is. Something that is a service AND an application — a git forge is consumed as a remote and operated through a web interface, and neither reading is wrong. Something that provides and consumes, because provider and consumer are ends of edges rather than kinds of module. Something the mesh installs that then becomes a node CAPABILITY, which means a node's provides-list is partly derived from what is installed on it and not only detected. Something that must be adopted rather than installed. Something that is a set rather than a thing. Something with exactly one instance for the whole mesh, where assigning it twice is not redundancy but two meshes. Something that is not software at all — a firewall policy, a DNS record, pure desired state, which fits the host's declaration model exactly and an installable package model not at all. An agent. The host itself, which is not a module and needs a schema that can say so. And the things the mesh depends on and does not control, which are why a node can be perfectly configured and still not work. Nine axes come out of it. Two are covered by nothing anyone has proposed: HOW MANY INSTANCES a thing may have, and WHETHER TWO CAN COEXIST — `excludes` covers part of the second and nothing covers the first. And one question the cases sharpen: is "runs" a property or a kind? The axes say property — one schema with a field saying how it runs, `never` included. The alternative is several kinds of module with different schemas, which is the taxonomy this effort already rejected once for services and applications. |
||
|
|
f55ecc1a47 |
011: an abstract name needs providers that are actually substitutable
Two corrections from the operator, and the first improves the design rather than narrowing it. `database` is not an edge. The test it fails, and the test the proposal was missing: can a consumer be switched from one provider to another WITHOUT CHANGING? A module speaking Postgres does not speak MongoDB or SQL Server — different wire protocol, dialect, driver — so a consumer declaring `requires: database` and handed any of them breaks. The name promises what no provider can deliver, and the resolver would report a requirement satisfied that is not. `terminal` passes: anything that runs a command in a terminal works and the consumer never learns which it got. So the ADAPTER is what creates an interface. `ai-assistant` is legitimate exactly because adapters normalise what is behind it. Without one there is no interface, there is a category — and a category is a TAG. Tags describe, edges bind, and keeping them apart is what stops the catalogue acquiring a second kind of relationship that looks like a dependency and is not, which is what a folder named after a domain already was. And the domain module goes. A `networking` module gathering a firewall, a resolver and a proxy under one name came from an older shape and does not fit — there is no such thing to install. There is core infrastructure: concrete modules named individually, not flavourable, with no grouping module standing in front of them. Fixed three places where the revision left the old rule standing, including an example manifest still requiring `database` — the kind of contradiction that would have been read as the design rather than as a leftover. |
||
|
|
e20a09ae80 |
011: one kind of edge
The design, rather than an account of what exists. A module provides names and requires names, and that single relation absorbs three things this effort had listed separately: requiring another module is requiring a concrete name, requiring a resource is requiring an abstract one, and an interface is simply a name with more than one provider. Nothing has to declare that it is an interface — it either has one provider or several. The move that does the most work: a NODE provides names too. Its profile is a set of them — display-server, container-runtime, an architecture — so a module requiring a display server is satisfied by the node exactly as one requiring a database is satisfied by another module. One resolution instead of two, and a graphical application cannot land on a node without a display server for the same reason, through the same code, that it cannot land without its libraries. Which makes the host's capability detection an input to resolution rather than something a person reads. It was built to be read; it turns out to be a provides-list. `excludes` is the one genuinely new relation, because it is not derivable: two modules that both provide message-bus look interchangeable when installing both would break the machine. Constraints are not placement. They say what must be true of a node, never which node — which is the mistake the measurement found in the current catalogue, where a module pins its database to a named node so a second node cannot provide it without editing the consumer. What it deletes, for the design: the module/resource distinction, the interface as a kind of thing, capability checking as a separate mechanism, domain grouping — folders assert relationships where edges record them, so a domain becomes a query over the graph rather than a directory somebody keeps true — and possibly tiers, if a tier is just a computed level. What it does not delete, stated so it is not discovered later: a resolver still has to exist, with version constraints and conflicts, and the design owes an answer on what it delegates rather than reimplements. |
||
|
|
c3a2984b3e |
011: measured, and the premise was wrong — the graph is not missing
The effort was opened to ask whether the catalogue's missing structure is a graph. It is not missing. 126 manifests, 103 edges, no cycles, nothing dangling, deepest chain of five — and a resolver in the SDK that topologically sorts them, already called by the tool loader at startup, the installer when syncing modules onto a node, and the delivery coordinator when expanding what a change affects. It already does something this effort assumed would need designing: a requirement on another module's provision is treated as an implicit edge to the module that provides it. So "ordering by the graph", which ADR 0043 makes the control plane's job, is a thing to call rather than a thing to build. The one place the graph is wrong, it is wrong about the substrate. A module needing a database declares `provider: postgres` inside `provisions:` — which is what a module OFFERS — so the resolver, which reads `dependencies:` and `requires:`, never sees it. Three edges are invisible this way, and they are the mesh's own database, the mesh's own broker, and the work engine's database. The consequence is measurable: computing what a working mesh needs from the declared graph gives registry -> sdk -> mesh -> meshware. Four modules, four levels, no database. Arithmetically correct and obviously wrong, for exactly one reason — a field that means "depends on" is not read as one. That is 04-ISSUES/003 in a new form: not a key nothing reads, but a key read as something other than what it means. Two latent defects, both contrary to ADR 0008 and both in the component ADR 0043 makes responsible for ordering a host will apply without question: a cycle warns and falls back to input order, and a dependency that does not exist warns and continues. Neither has fired, because the catalogue currently has no cycles and nothing dangling, which is why nobody has noticed. And placement is decided in the catalogue: a provision pins itself to a named node in the manifest. Which node runs what is an inventory decision — tier 2 by the skeleton's own test — so a second node cannot provide the mesh's database without editing the module that consumes it. What the graph would DELETE is currently nothing. What it would add is three declarations that no manifest uses today: excludes, a required node capability, and an interface with adapters. Whether they would be used is not measured, and zero usage is equally consistent with nobody needing them and nobody being able to express them. |
||
|
|
278f7427ed |
012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly succeeded or failed, with a severity per line. Taken with one change — the overall is DERIVED as the worst mark present, never written alongside. Two fields maintained independently drift, and a briefing reading "full success" while carrying a failed line is exactly the fault this record keeps cataloguing. An outcome computed from its lines cannot disagree with them. Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success — adoption will meet configuration it cannot parse and state it cannot read, and folding those into "fine" is the same move as reporting an installed package as a capability. And adding severity reopens something the earlier rule did not cover. "Flags inform, they do not block" was decided about CONFLICTS, where the mesh chose deliberately and the machine still works. A failure is not "we chose" but "we could not". Treating both the same makes a node where something the mesh needed never happened indistinguishable from one where a log level differed. |
||
|
|
0106318bcb |
012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning for each is the useful part. What is already on the machine stays, the conflict is flagged, adoption completes. This buys non-destructiveness by construction: the class that made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — cannot arise, because nothing tied to the machine's physical reality is overwritten. It exposes the mirror. The mesh's configuration is not only preference; some of it is what a module needs to function. Keeping the machine's version there produces a module that is installed and does not work, which is 04-ISSUES/007 arriving from a direction that issue did not anticipate. And a fleet where every node kept its own settings is one where a module works on one node and fails on another with nothing able to say why. So neither direction is right as a blanket, and the question is not whose configuration wins. It is whether the module REQUIRES the setting or merely PREFERS it — required contradictions cannot be kept without breaking the module, preferences should always yield to what is there. That is a property of the module's declaration rather than of the adoption algorithm, which makes it one more thing the graph would carry. Until modules can say which of their settings are load-bearing, adoption is defaulting in the dark, and the default chosen is the one that does not break the machine it is adopting. |
||
|
|
bcb18c7329 |
012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's disagree, the mesh's version is installed, the conflict is flagged, and it is reconciled afterwards — the mesh's configuration is known to work, the machine's is not, and a half-adopted machine is a state nobody understands. So adoption always completes and flags inform rather than block, which also settles what 'adopted with open questions' prevents: nothing. The node is a node. The original is kept, so nothing is unrecoverable. One class left open rather than folded in, because it is the one place the oldest rule in this record argues the other way. 'Known to work' is true of the mesh's configuration in isolation, not on this machine. Most disagreements are preference and overwriting them is right. A few are tied to what is physically present — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — and installing ours there does not discard a preference, it can make existing data unreadable. Restoring the configuration file afterwards does not undo that. The default is settled. The exception is not 'there is a conflict' but 'applying ours would destroy something a configuration backup cannot restore', and identifying that class is open. |
||
|
|
60736199a4 |
012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort had open with two bad answers. Nothing is taken over without keeping what was there. Adoption happens on machines somebody is already using, and the configuration being taken over is configuration somebody chose. This is a never rule rather than a courtesy, and it earns that by the same incident the mesh's strongest rule carries: the worst loss in this record came from a tool acting on a path it did not own. Adoption is that act made deliberate, which makes the safeguard obligatory. And adoption produces a briefing, not just a result. It meets things a script cannot decide — a runtime configured one way against a mesh wanting another, a package pinned for a reason, local settings the mesh has no opinion about. Silently winning is wrong in both directions and refusing outright makes a machine in use unadoptable. So conflicts are FLAGGED: what it found, what it took over, what it could not resolve, written to be read by a person or an agent as the first thing a session on that node has to work with. That is the declaration parser's principle at a larger scale — name every problem at once, to somebody who can act on it. The question it turns on is recorded rather than assumed away: are flags advisory or blocking? A briefing nobody opens is worse than a failure, because the machine is in service and the record says it went well — 04-ISSUES/003 again. Working position: the node is usable and the mesh KNOWS it has unresolved adoption questions, as a state something can ask about rather than a document in a log directory. What that state prevents is undecided. |
||
|
|
ddb8091f68 |
Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not. The host can be told to run a container or install a package; both need a file, and asking where the host gets it produced a bad trilemma — carry everything, download at apply time, or push the files in first. Downloading fails on the first node, which cannot fetch the image registry from the image registry it is trying to start. The reframing came from the operator: the machine is not offline, and what matters is WHEN the fetching happens. Move it from apply time to build time — build the installer on a machine with a network, tailored to the target, apply it on a target that then needs nothing. The same move the lab already made for its router image. Which makes the question not where artifacts come from but what is missing from THIS machine, and that needs two things answered: the closure for a one-node mesh, and how a machine already in use becomes one. Adoption is the second half, and it is sharper than it sounds. Having a package installed is not owning it: a container runtime found already present carries settings somebody chose, and noticing the binary exists discovers none of them. It was also the original path — 00-as-is/05 records adoption of a pre-existing machine's configuration as the original mechanism, since made legacy and explicitly out of scope for the lab. It returns for a different reason than it was dropped for. Two collisions recorded rather than discovered later. ADR 0004 has managed files generated and never edited, and adoption needs a one-time import before that rule starts applying — three states, and the middle one is new. And ADR 0043 says the host never touches what it did not create, which is exactly what adoption does; that rule needs a companion rather than an exception. Eight open questions, including whether 'tier' is just a coarse view of a graph level, whether owning a package means owning its version, and what cannot be precomputed at all — because tailoring moves the cost of building from source rather than removing it. |
||
|
|
64913ed0d3 | Merge pull request 'The approval is the checkpoint, and what a declaration is' (#10) from design/approval-is-the-checkpoint into main | ||
|
|
00d5ba8376 |
0042 and 0043, as approved
0042 becomes the operator's own rule and stops there: every merge into the main branch is notified and approved. Notified means proposed and said out loud, not performed and mentioned; approved means a person says yes to THAT merge. Who performs it is not the thing worth constraining, which is what makes an agent merging its own work unremarkable — the checkpoint already happened. 0043 answers the question it did not: where the ordered list comes from. By hand today, in substrate.lock, because the first node has no control plane to derive anything from. Afterwards the control plane derives it from module assignments, resolved configuration, and what each module declares it needs — ordered by the dependency graph, which is research 011. So the record is complete on the consumer side and deliberately silent on the producer side, and that is a legitimate order to settle them in: the host must refuse what it does not understand whoever wrote it. One consequence that only appeared when the question was asked: if the graph turns out not to determine a total order, that is 011's problem and not the host's. The host is still handed a list and still applies it as given. Recorded because it is the seam where a future difficulty would otherwise try to migrate into tier 0. Both were marked accepted before they had been read. Approved now, so the field is true — which it was not when it was written. |
||
|
|
bea052753e |
ADR 0043 — what a declaration is
Stage 2 could not start without it. Three constraints already bound the shape and between them they decide most of it. JSON, because the standard library carries it and carries no YAML, and a YAML declaration would put a third-party parser inside the one binary whose whole argument is that it needs nothing — to gain authoring comfort in a document generated by a machine and read by a machine. An ordered list, because ordering is a DECISION. A host deriving order from declared dependencies would be deciding the thing most likely to differ between what the control plane intended and what the machine does. The control plane knows what depends on what; it says so by saying when. Unknown is refused, never skipped — an unknown version, type or field refuses the whole declaration. A host that skipped what it did not understand would apply most of a declaration and report success, which is 04-ISSUES/003 with the declaration on the other side of the wire. Complete for what the host OWNS, and only that. It removes what it previously applied and is no longer declared, which it knows from the store rather than by inference, and never removes what it did not create — a converger that treats 'not declared' as 'must not exist' deletes what the mesh never put there. Two consequences arriving earlier than the build order suggested: the store is load-bearing at stage 2, because nothing can be removed without knowing what was applied. And a closed address space bounds the first vocabulary to what needs no network, because a scenario has no route to a package repository. |
||
|
|
23232e019a |
ADR 0042 — the approval is the checkpoint, not the second pair of hands
§2 said "never merge your own", written for people. Applied to an agent it produced a contradiction that surfaced immediately: an agent asked to merge cannot merge, because it authored what it is being asked to merge. So every merge here was either performed by the thing that wrote it, or not performed. Rejected the literal reading, because an operator clicking merge dozens of times without reading is not a checkpoint — it is the SHAPE of one, which is worse, since the record then claims a review that did not happen. Rejected dropping the rule, because the failure it prevents is not one an agent is less prone to. So the rule names what the checkpoint actually is: a person deciding, not a person clicking. Work may be merged by whoever wrote it once a human has explicitly approved that merge. What "explicit" excludes is the half that can rot, so it is enumerated: a standing permission cited forever, an instruction to do the work read as approval to merge it, silence, and the author's own judgement that it is ready. This narrows rather than relaxes. The obligation moves from who performs the merge to whether a person decided — a higher bar in the case the old wording permits, where a reviewer merges someone else's work without reading it. Synced, and verified by reading the rule back out of the live page rather than by trusting the publish. |
||
|
|
7ef1c561e0 | Merge pull request 'Tier 0: the questions answered, the decisions taken, and the design' (#9) from design/close-the-record into main | ||
|
|
92e8c74ce4 |
ADR 0041 and the build handoff for the node host
Building tier 0 forced the question "the one binary installed by hand" had been carrying unexamined. A TypeScript host needs a runtime present before it runs, so the thing installed by hand becomes two — and the second must be installed by the means the host exists to replace. So the host is a statically linked binary that requires nothing present, written in Go. Rejected: a runtime installed first, which breaks the property the tier rests on; and bundling the runtime into the executable, which carries ninety megabytes to preserve a language choice and puts a young feature at the bottom of the stack. The argument that decided it is architectural rather than about taste. 0037 means the host never queries the mesh database and 0039 means it only receives declarations, so the host shares NO code with any other tier — not a client, not a schema, not the SDK. The language boundary falls exactly on a boundary that already exists, and a second language usually costs duplicated logic where here there is none to duplicate. §8 gains a scope: it said "TypeScript throughout" when everything was a service or a surface, and is now scoped to those with tier 0 named. Another sync owed. Playbook 04 steps 2 and 4: repos.md records mesh-host as existing, the design takes code: [mesh-host] and status: in-progress. |
||
|
|
b9facf9375 |
Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply declared state on this machine — and the six absorbed concerns are instances of it, not additions to it. Specifies the six parts and what each owns, and the two properties that make apply trustworthy rather than merely present: every applier reads back, because setting a value is not evidence the value took; and what was applied is recorded after it works, never before, because a failed apply leaves the machine wherever it reached and nothing must claim otherwise. Build order is staged so each stage is verifiable in the lab before the next exists. Stage 1 is profile and inventory — no control plane, no declarations, no network — and it is deliberately the smallest useful thing, because `place:` has nothing to place and the lab therefore raises empty machines. Stage 1 ends that, and every later stage is tested by a lab that already works. Stage 2 is the one that could invalidate the tier boundary: whether one host can raise the substrate alone is Move 1's assumption and has never been proved. Every decision the design rests on is given the test that asserts it, per 0034 — including the dependency-direction lint, which is what makes "the host never queries the mesh database" a rule rather than an intention. Six things left open and named, including the one that host-size.md could not measure: zero dependencies, but still six vocabularies. |
||
|
|
6e4fc5d69b |
ADR 0040 and the constitution sync — absorb, then publish
The sync came due for three accepted rules. Reading the target before overwriting it found the source and the enforced copy had diverged unrecorded, and that a literal republish would have DELETED rules the mesh enforces: the live page's §4 carried SOLID, layering, TypeScript strictness and DRY/YAGNI, which appear in no decision record anywhere and have been checked against for six weeks. Removal was not the safe alternative either. The orchestrator reads "when absent, no constitution is injected (backward-compatible)" — so deleting the page would not fail, it would silently inject nothing, and every design meeting would run unchecked. Three unenforced rules would have become all of them. So: absorb first. The code-quality rules land as §8 rather than §4, because appending renumbers nothing and every existing citation stays valid. They are marked as inherited — every other rule states the incident behind it, these state nothing because nothing was written down, and importing them silently would have claimed a provenance the document does not have. The review bar is resolved to a person who is not the proposer. The live page required two node operators; there is one, so the rule was never met and could not be — a rule that cannot be satisfied is not a high standard, it is one everything silently violates. Then published, and the read-back earned its place in the playbook: the FIRST publish reported success and changed nothing. New revision, new title, body unapplied — a malformed argument dropped silently. §5 demonstrating itself during its own publication. |
||
|
|
34e7a4780a |
Accept 0035 — a report is read from the system
The principle is obvious; the obligations it imposes are not, and those are what was actually being decided. The applier records what it did, including facts it never reads itself, purely so something else can read them back. It records them AFTER the thing works, so a half-finished run ends up with less metadata rather than optimistic metadata — rejecting the simpler alternative of tagging at creation with a status field, because a failed raise deliberately leaves wreckage standing and wreckage tagged at creation claims things that never happened. And it binds anything that reports on the mesh, not only the drawing. §5 carries the rule unqualified now. The constitution sync is owed for this and for 0034 and 0018, and has not been run. |
||
|
|
66a89cefbe |
Accept 0039 — a node owns no password, only an identity
Settled in the operator's own words: nodes should not own passwords, only an identity when communicating to the broker. Recorded that way at the top of the record, because it is the whole decision in one line and the rest is why. What this obliges, in order of newness: enrolment is the one mechanism that does not exist. Per-node broker users, virtual hosts and per-queue permissions are broker configuration. Mutual authority is certificates on a connection already open. And the boundary must fail legibly, which is the requirement the debugging objection earned. |
||
|
|
9ee0c63c7a |
0039: the link already exists, and most of the cost is already paid
Written first as though the link were a thing to build. It is not. ADR 0001 already has it — every node connects outbound to a single broker, nothing ever connects to a node, each node declaring an exchange named for itself and consuming from its own queue. Already outbound-only, already per-node addressed, already the one channel everything arrives through. So this record is not proposing a channel. It proposes that the channel carry per-node identity instead of one shared credential. The same as-is records the fault it fixes, for the broker rather than the database: the broker's credential is mesh-wide, rotating it is a mesh-wide operation, and doing it wrong has taken the broker down. That reframes the overhead objection, which was fair against what the record said and not against what it means. Of six properties, four are already true, one is broker configuration — users, vhosts and per-queue permissions the broker already implements — and exactly one is new machinery: enrolment. Meanwhile 0037 subtracts, since a node under it holds no database credential at all. Three hand-carried shared secrets become one identity that grants only identity. Adds the option that was actually being weighed and was missing: accept the exposure as the cost of simplicity. Rejected because the simplicity IS the unfixability — the credential cannot be rotated precisely because everything holds the same one. And adds a requirement from the debugging objection, which was the strongest part of it: it must fail legibly. A boundary that refuses a node without saying why is worse than the password it replaced, because a wrong password at least announces itself. That is §5 applied to a security mechanism. |
||
|
|
902739acb6 |
Research 011 — the module graph
The proposal to split modules into provisioning services and applications was worked through and abandoned, for a reason worth keeping: it cannot be filed consistently. A git forge is consumed as a service and operated through a web interface; an analytics service grants tracking identity and is a dashboard. The operator's correction is the sharper form — what runs on the machine is a supervised container, not something a user started. That is a fact about HOW a thing runs, not about what kind of thing it is. So it is a facet, and 0002 survives: everything is a module. What the catalogue is missing is not a taxonomy but a graph. Grouping asserts relationships; a graph records them. Five declarations, of which two exist: requires/provides a resource (yes), requires/excludes another module (no), requires a node capability (no). Plus interface modules that carry no implementation, with adapters providing them. Recorded because it matters: this is a package manager's model, and pacman already has all of it — depends, conflicts, and provides as virtual packages, which is exactly the interface/adapter idea. Arriving there independently is evidence for the shape. It is also a warning about what not to reimplement. Working position on capabilities, to be tested: intrinsic ones (hardware, architecture, network position) are detected and never installed, and a module requiring one it lacks is impossible rather than unresolved. Provided ones (a display server, a container runtime) are not a separate kind of thing — they are modules that provide a capability, so "may the mesh install a capability" is not policy, it is dependency resolution. Issue 007 then bears directly: an installed package is not a capability. Also captured: the operator's assessment that the machinery around a module — scheduled tasks, hooks, migrations, config and env — is worth keeping, seeds are not, and the integration is wrong enough to need a major refactor. Research 005 found supporting evidence from another direction, that the densest apparent coupling in the catalogue is manifest boilerplate churn. The first open question is the one that decides whether this is progress: what does the graph DELETE? If modules gain declarations and lose nothing, it is motion. |
||
|
|
f5073b9b00 |
Accept 0034 and 0018; 0011 is superseded
0034: a test defends a decision. §5 carried it marked "proposed, pending review"; the marker is removed and the rule now stands unqualified. The lab was already built to it, which is the inversion §5 exists to catch — closed now rather than left standing. 0018: the mesh creates no symlinks. §3 said the installer owns the links today and the INTENT was that the mesh creates none. It is no longer an intent, so the wording says so, and ADR 0011 becomes superseded rather than edited — its reasoning is why the rule exists at all, and the incident behind it is the reason anyone believes either record. The links the installer still reconciles are a migration, not a permission. The constitution sync (§6 step 4, playbook 05) is NOT done. An unsynced rule is a rule the mesh does not enforce, whatever this document says — and publishing it changes what every design meeting is checked against, so it wants saying out loud rather than doing quietly. |
||
|
|
4866007c04 |
ADR 0038 accepted; ADR 0039 (proposed) — the link is the security boundary
0038 accepted: no two modes. One behaviour, two sources of declaration. 0039 fills the gap 0038 named. Proposed rather than measured — it designs a boundary that does not exist — but what it replaces IS measured, and that is the argument. Today adopt.sh asks the operator to paste in the postgres password and the object store password, the same ones on every node, and they are not discarded after adoption: wireguard and traefik open a pg connection on every reconcile. So every node permanently holds a credential to the control plane's database, and 00-as-is/06 records that nothing rotates it. Compromise of any node is compromise of the mesh's store, with no way back. Four properties make the link a boundary rather than a pipe: it is outbound and node-initiated, so a node has no listening control surface — which the topology already requires, since most nodes have no forwarded port. A node holds its own identity and nothing else, so compromise of a node is compromise of that node. Authority is mutual, because a host that applies whatever the link delivers must know the mesh from something impersonating it. And what may be pushed is bounded by FORM — declarations of known shape, never a command to run. That last property is stated with its limit rather than oversold: it bounds form, not impact. A compromised control plane can declare harmful state and the host will apply it faithfully. What it buys is a describable blast radius. Joining uses a one-time short-lived enrolment token, useless once used and useless after a while, in place of hand-carried shared secrets. Named rather than hidden: the first node's identity is self-issued and becomes the root of trust, which is the one place 0038's "no special first node" does not fully hold. Rotation becomes possible and is still not designed. And what may expire is constrained by 0036 — an identity needing refresh would make a laptop fail for being a laptop. |
||
|
|
72b22830f3 |
ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation. The open question posed a class distinction — full nodes and lesser presences. There is none. Reachability is state, not kind, which promotes the host's local store from a component to a requirement: it is what makes disconnection ordinary rather than exceptional. The reduced contract the question reached for is real but it is capability, and that belongs in the profile. 0037 (accepted): the host applies, it does not decide. Measured rather than argued — the absorption is smaller than the machinery that already applies state, and eight of ten adapters carry no dependency to move. The two that do open a Postgres connection to the control plane, which inside tier 0 is the one thing the tier rule exists to forbid. So each concern splits: deciding needs every other node and stays in tier 2; applying needs root and locality and goes to tier 0. The host carries ONE concern, of which the six are instances. 0038 (proposed): a node joins by linking first. The operator's two-modes proposal, adopted as intent and corrected as structure. Two modes is two code paths where the first runs once per mesh and rots — and the mesh already has that fault in its worst form, as three hand-run shell scripts. Instead: one behaviour, two sources of declaration. The first node is not a different kind of node, it is a node whose mesh is not up yet, and its specialness is temporary and self-erasing. 0038 also shrinks the migration 0037 called expensive: a joining node never needs mesh-wide state, because the hard part of the overlay is only needed to compute the WHOLE mesh. It needs one peer. The rest arrives. Left open and said so: what may be pushed over the link and how a joining node proves it is entitled to join, and whether one host can raise the substrate alone. |
||
|
|
42bce02bba |
006: answer the host-size question by measuring it
The skeleton's biggest unproven claim was that absorbing six concerns makes a binary whose whole argument is having no dependencies carry six of them. Measured against origin/main, and the question turns out to ask about the wrong axis. By size the absorption is SMALLER than the machinery that already applies state on a node — 2755 lines of adapters against 3059 lines of meshware, env-sync and config-sync. The host is not a new large thing; it already exists, spread across three core modules. The real risk is direction, and it is two modules wide rather than six concerns wide. Eight of ten adapters already receive derived state and only apply it, so absorbing them moves code that has no dependency to move. Two — wireguard and traefik — open a Postgres connection to the control plane and compute their own configuration, which inside tier 0 would be an upward dependency and is exactly what the tier rule forbids. And the split has already been happening without being named: dnsmasq-app needs the same node data as wireguard and does not query for it, because hand- duplicated state went wrong and someone derived it centrally instead. Eight of ten adapters are on the far side of that migration. So the absorption is not a move, it is a split: deciding stays in tier 2, applying goes to tier 0. The claim survives with its scope corrected — the host carries ONE concern, apply declared state on this machine, of which the six are instances. Stated open rather than glossed: the two unsplit modules are the two hardest, six concerns is still six vocabularies even at zero dependencies, and what the host must carry versus find is issue 007 and unresolved. Question B also recorded as answered by the operator — a node is a managed machine, and a disconnected node is still a node in a different situation. The question posed a class distinction; there is none, and what varies is state. |
||
|
|
4bf7a35568 |
Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design, design before build — and skipped for the bookkeeping. This closes that. 004 graduates. Its one open item was "not yet stood up"; the lab is stood up, and the substitution the effort turned on is now enforced by the validator before anything is raised rather than left as a thing to remember. Its certificate conclusion has a home in 01-end-to-end-testing and is designed but not built — implementation is a third axis, and an effort graduates on its conclusions. One item leaves 004 without a home and is recorded rather than lost: the reverse proxy does not set caServer, so it defaults to the production endpoint. The two lab designs read `designed` while running in production of a sort, so they become `in-progress`. And the lab gets an as-is document, which it did not have. It records what runs including the parts nobody would choose again: that `place:` is refused and the lab therefore raises EMPTY MACHINES, that the drawing shipped with no design document behind it, that a router is tagged as a machine for a reason found by a bug, and that the integration suite raises two of five scenarios while both faults found so far lived in the three it does not. 006 stays active, deliberately. Two of its open questions ARE the tier 0 design — whether absorbing six concerns makes the host too large, and whether an unprivileged node earns a place in the inventory. Playbook 04 is explicit that an open question is a reason to research, not to build around. |
||
|
|
b225b07625 |
ADR 0035 (proposed): a picture is read from what runs
Drawing a scenario forced a choice that looks cosmetic and is not. A diagram built from the declaration and captioned "as raised" answers "is what is running what I asked for?" with the request, which always agrees with itself. So: a picture captioned as raised reads only the running system, and where the hypervisor does not hold a fact the picture needs, the raise records it on the resource. With the rule that makes the recording worth anything — a behavioural tag is written after the behaviour works, never at creation, because a failed raise leaves wreckage standing and a picture of wreckage must not badge what the wreckage was supposed to be. It earned itself on the first comparison: every VM showed no addresses, because a container's interface carries the device's name and a VM names its own. The two pictures disagreed, so a whole class of machine silently losing its addresses was visible in seconds. §5 of how-we-build gains the general form, marked proposed. The constitution sync is deliberately NOT done — a rule the mesh enforces before a second person agreed to it is what §6 exists to prevent. |
||
|
|
8efa063f21 |
The snapshot question is answered by a test
The lifecycle design asked whether a scenario snapshot needs the machines stopped. The integration test answered it on its first run: no, but they must be flushed. A snapshot captures disk and not memory, so a write still in the guest's page cache is absent from it — not stale, absent. A file written seconds before a snapshot did not survive the restore. Flushing first buys write-durability. It does not buy application-consistency: anything mid-transaction is still captured mid-transaction, and that limit is now stated rather than left implied. |
||
|
|
a2495e4d8e |
ADR 0034 (proposed): a test defends a decision
how-we-build 5 already says that if a document states a rule about the mesh, it says how the rule is verified — an unenforced rule being indistinguishable from a wrong one, and costing more because people believe it. That has never been applied to decisions, and a decision record states the same kind of claim. The gap was found by review: the lab reached 2,128 lines with 1,072 untested and no stated rule broken, because there is no testing posture in how-we-build at all. Every decision the lab embodies was verified by hand and none of it survives the terminal it ran in — which is 04-ISSUES/005 in miniature, coverage assumed rather than checked. Rejected a coverage percentage: it measures how much code a test touched, not whether anything important is defended, and would have been satisfied by testing the parser harder while the hypervisor integration stayed unasserted. Rejected test-driven development as a hard rule, and not because it is wrong in general. Half this implementation was discovery — that the hypervisor CLI reads a definition from stdin and hangs, that it assigns a MAC without recording it, that a stock image's networking flushes a static address. A test written first against undiscovered behaviour asserts a guess. So: structure and logic tested first, behaviour against a real system tested alongside, mocking the boundary forbidden, and a blocking gate as the definition of done. A test names the decision it defends, which is what makes the pairing checkable — a decision without one can be found rather than noticed. Stated as proposed rather than adopted: 6 requires review by someone who is not the proposer. Records 0001-0033 predate it and are not retroactively invalid, but each should acquire a test or an explicit note that it cannot have one, and until then the rule is aspirational for them — which is the state 5 warns about, recorded rather than hidden. |
||
|
|
eab4598494 |
ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is fidelity: a node boots a stock image and runs the real install, so it has to be a real machine or the thing under test is not the thing that ships. That reasoning does not reach a router. Nothing under test runs on one, it holds no identity, the mesh never installs anything on it, and no assertion is ever made about its internals. It exists so packets behave the way they behave in the world, which is the definition of scenery. So a router is a system container. What it must reproduce is kernel behaviour — translation, connection tracking, filtering, forwarding — and a container has the same kernel. Verified before deciding rather than assumed. In a plain unprivileged container: ip_forward and ipv6 forwarding both settable, nftables masquerade accepted and listed back, and the conntrack timeouts that mapping_ttl depends on both writable. No privileged mode, no nesting, no capability grants. Rejected letting the hypervisor provide NAT, on a stronger ground than speed: it makes the lab provide what the declaration is supposed to own, and it cannot express a mapping that expires, a gateway that refuses to forward, or policy between siblings. The model would shrink to fit the tool. The distinction is now load-bearing and has to stay legible: node means something under test, scenery means something that makes the test real. If the mesh ever installs anything on a router, it has become a node and this record no longer covers it. |
||
|
|
132f6a0626 | Merge pull request 'The scenario model, the lifecycle, and what the lab actually costs' (#6) from design/scenario-underlay-detail into main | ||
|
|
e88b448145 |
The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing 1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle falls from ~90s, unbounded at worst, to ~15s dominated by a boot that cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write and did not without it. The consistency matters as much as the speed: three consecutive snapshots took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never finished. One honest counter-observation recorded: launching onto the fresh copy-on-write pool was slower, 20.2s against 14.3s, because the image had to be unpacked into a pool that had never seen it. Paid once per pool, and dwarfed by what snapshotting saves, but it went the other way. Doing the measurement produced the answer to how the lab installs on a clean machine, because both failure modes appeared while doing it. Installed is not available: the daemon was present with units disabled and no group. Issue 007. Available is not adequate, and this is worse: with the storage tooling absent everything worked and snapshots were seventy-six times slower. Nothing failed, nothing warned. That is a variant the mesh has not catalogued — its usual failure is reported success and did nothing; this is reported success and did it seventy-six times slower, which no error surface catches because nothing is wrong. So the lab verifies CAPABILITY, never installation, and refuses to run degraded rather than warning — a warning about a slow inner loop is read once and ignored forever. Prerequisites may arrive from a mesh module or from the lab's own bootstrap, and the second path is required rather than convenient: a lab installable only by a mesh cannot host the development of the mesh that installs it. The lab is the second thing installed by hand, after the node host, and for the same reason: something has to be first, and pretending otherwise produces a circularity papered over by a script nobody exercises. |
||
|
|
98bcd5cc49 |
Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than arguable, so it was measured. One virtual machine on a workstation with hardware virtualisation and NVMe. Raising: the launch call returns in 3.4s, the machine is actually usable after 14.3s. The gap is a design constraint — raise must wait for the second number, because reporting the first would be transport reported as effect, which is the mesh's own recurring failure. Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full copy; nothing is shared. Restore: 10.4s, usable again after 20.1s. The second snapshot exceeded two minutes and never completed. That is the more troubling number: snapshot cost here is not merely high, it is unpredictable, and a loop with a variable multi-minute step is one nobody trusts. Projected to a four-machine scenario, a reset-and-rerun cycle is about a minute and a half at best and unbounded at worst, before any of the mesh's own work begins. That is too slow for an inner loop, and ADR 0029's whole argument — that making the bootstrap path the inner loop turns the least-exercised code into the most-exercised — holds only while resetting is cheap. The cause is not virtual machines. Hardware virtualisation is present and machines boot in fourteen seconds. It is that the daemon offers exactly one storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The btrfs kernel module is available; btrfs-progs is simply not installed, which is the entire reason the driver is absent. The copy-on-write comparison was deliberately NOT run, because running it would mean installing a package by hand — which the rules forbid and which would have made the measurement unreproducible. So the honest statement is that the current configuration is too slow and the likely fix is known but unverified, rather than that btrfs fixes it. |
||
|
|
a253afe020 |
Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. |
||
|
|
e88a6df924 |
Public networks are unrelated, and routed rather than bridged
Caught in review: every public address sat in one /24, which made the three of them look like one network. They are not. The internet is a very large number of unrelated networks routing to each other, and a machine in one is many hops from a machine in another with no shared broadcast domain between them. Putting them in one prefix would have quietly made four false things true in the lab: machines resolving each other by ARP and talking directly, TTL never decrementing, broadcast and multicast crossing between them, and any two being adjacent. The third is not hypothetical. The as-is layer records that mesh names are deliberately not multicast names, after a delay and a one-node-only failure mode. A lab where the internet is one broadcast domain would let a node discover a peer by multicast that it could never discover in production, and report success — the exact false green this effort exists to prevent. So a scenario has one public segment per public NETWORK, each with its own unrelated prefix, wired together through a router and never onto a shared bridge. That is a property of how the lab wires them rather than a field anyone sets, because no correct scenario has two public networks adjacent. Addresses now spread across all three RFC 5737 ranges plus RFC 3849 /48s, chosen to look nothing like each other, and a foreign private network uses someone else's RFC 1918 range rather than a documentation one. All three examples in the document rewritten, since two of them still showed a single flat internet segment and contradicted the new rule. |
||
|
|
a873088140 |
A worked example: the whole model applied to an ordinary mesh
The shape research 004 identified — one machine with a routable address, one publicly named but behind a household connection, one stationary on that network, one that roams — written out with every field the model has, in role names and documentation addresses. It shows the ISP modem doing nothing, because in bridge mode it is a media converter: it changes the physical medium and leaves the packets alone, so it creates no IP-level fact and appears nowhere. In router mode it would be a second gateway and publishing would need a rule on both, which is the one case the model still cannot express. It shows two segments sharing one gateway declaration, which means one gateway machine, and a policy rule between them that is asymmetric because useful ones almost always are. And it shows what is deliberately absent. Research 004 recorded overlay addresses, hub election and names for exactly this topology, and none of them appear: a scenario must not state what the mesh is responsible for. Given the declaration, whether a hub is elected, whether the NATed machine's endpoint is learned, and whether the roaming machine re-forms after moving are all observed rather than arranged. The absence is the point. The run at the end moves one identity through four positions — home, foreign network, asleep, home again — against a foreign gateway whose mapping expires in 30 seconds, which is why a number is there rather than a boolean. |
||
|
|
b1874f1d0f |
Segment policy, shared gateways, and what the model leaves out
Asked whether a real setup is coverable — router, modem, access points — the answer splits, and one part was a genuine gap. Most equipment is invisible and the omission is deliberate. The test: does the device change what an IP packet can do? A switch moves frames within a segment. An access point bridges wireless clients onto one — a machine on wifi and a machine on cable are the same machine to IP. A controller configures equipment and has no packets of its own. Modelling any of them adds a fixture with no fault to catch. Two entries in that list do matter. A modem in bridge mode is a media converter and invisible; in router mode it is a second gateway, which is double NAT — expressible as nested segments, but publishing through two gateways still is not, and that is now named as the one real absence. And VLANs are segments, which exposed the gap: inter-segment policy was inexpressible. inbound: is a HOST firewall, per machine. A segmented router enforcing rules between networks is a different thing and blocks traffic regardless of what the destination thinks — a node behind such a rule cannot be reached even by a peer that knows exactly where it is. policy: states it as a fact about a pair rather than a property of either, defaulting to allowed and asymmetric by design, because the useful configuration is almost always one-directional. Segments may also share a gateway: identical gateway declarations mean one gateway machine, not two, because that is what a VLAN-capable router is — and two routers sharing an address would not work anyway. |
||
|
|
274bd3b304 |
Close the missing axes — and address family changes the model
Address family was not a field. IPv6 usually has no NAT, so a machine behind a household gateway is typically unforwardable on v4 and DIRECTLY ATTACHED on v6, at the same moment. The three positions therefore apply per family, and reachability is a property of (machine, family) rather than of a machine. The consequence is bigger than the syntax: 'can these two nodes reach each other' stops being a yes/no question. It is asked once per family, and the asymmetric answers are the interesting ones. A mesh treating reachability as one fact per node reaches a peer over one family, fails over the other, and reports whichever it tried. That distinction did not exist in the model and would have been found by a failure rather than by reading. Two fields follow from it. inbound: allow|deny became necessary because with NAT unreachability was implied by topology, while a globally routable v6 address is reachable unless something refuses — so refusing has to be sayable or v6 addressing silently implies reachability. And nat: became a list of families rather than a boolean, because a real gateway translates v4 and routes v6 and a boolean cannot say that. mapping_ttl closes the keepalive gap: a mesh holding a connection through NAT without refreshing it works perfectly until the far side goes quiet for longer than the mapping lives. segments[].mtu closes the fragmentation gap: an overlay adds a header, so a tunnel over a reduced-MTU path establishes a connection and then silently drops large packets. at: takes a list, so a multi-homed machine is expressible — which the model already implicitly required, since a border machine sits on two segments. v6 uses RFC 3849 documentation space, the exact counterpart of the RFC 5737 rule and load-bearing for the same reason. Remaining: nested forwarding and an address changing in place, both extensible when needed. Path quality stays deliberately out — it changes performance, not correctness, and modelling it makes a network simulator rather than a fixture. |
||
|
|
b944904f1a |
Audit the scenario model for generality, and fix what it found
The question is not whether the model covers our mesh but whether it can express any mesh. Audited against the axes a deployment varies along, with the standard being every property that changes how the mesh BEHAVES rather than every property a network has — bandwidth does not change correctness, MTU does. One real bug, now fixed. A segment with no gateway was read as the internet, which made an isolated network inexpressible: a LAN with no route out would have been treated as public and forced onto documentation addresses. Segments now state kind: public or private, and a private segment with no gateway is an island. A mesh spanning a site with no internet is a real topology. One modelling error, now corrected. The three positions were framed by ownership — a gateway you control versus one you do not. The axis is forwardability. Carrier-grade NAT is your own connection and is still unforwardable, so it belongs with the café network. Gateways gain forwardable:, independent of nat:, and publishing through an unforwardable one is a declaration error because that is the constraint being reproduced. Three genuine gaps recorded in priority order. Address family: cidr is implicitly v4, and a v6-only node is not exotic — a mesh that assumes v4 fails there completely rather than partially, which makes this a second world rather than a refinement. Expiring NAT mappings: without them keepalive behaviour is hoped for rather than tested, and for a mesh mostly behind NAT that is the fault that shows up after an idle night. MTU: tunnels fragment, and a smaller-MTU path establishes a connection that then silently drops large packets — the exact shape this effort exists to stop shipping. Latency and loss are deliberately out: they change performance, not correctness, and modelling them makes a network simulator rather than a fixture. Also adds a NAT primer, because the three positions are consequences of it and the document should not assume the reader already knows why a mesh dials outward and never inward. |
||
|
|
e65e5809dc |
The scenario declaration gets a real network model
forwarded: [443] was the tell. It implied a destination-NAT rule while never saying from which address, and the address is the whole point: a household's public address is what a peer records as the endpoint when a machine there dials out, and what a public name for a published machine there resolves to. It was decoration in the old shape and is load-bearing in this one. The model now names three positions a machine can be in, because they are genuinely different and the mesh has to cope with all three. Directly attached, with its own routable address. Behind a gateway you control, reachable only through a forwarded port at the gateway's address. Behind a gateway you do not control, reachable not at all, with an apparent address belonging to someone else's router that changes when the machine moves. The third is the hard one and the one that breaks reachability assumptions first. A gateway now carries three facts instead of a boolean: the parent segment, the address the world sees the network as, and whether addresses are translated — so a routed range is expressible as well as ordinary household NAT. published names the gateway it forwards through, which is how a machine on a LAN that itself has a public address is stated, and publishing on a foreign gateway is a declaration error because that is exactly the constraint being reproduced. Moving a machine between positions becomes a lifecycle operation rather than a declaration: the same identity at home, then on a foreign network, then asleep, in one run. Whether the overlay survives that and notices the endpoint changed is observed, never arranged. All three RFC 5737 ranges are now allocated a job — the internet segment, a foreign network, and a spare — with private segments kept byte-identical to production because those addresses mean the same everywhere. New open question worth having: a real gateway forgets NAT mappings after a timeout, and whether a scenario can say so decides whether keepalive behaviour is testable or merely hoped for. |
||
|
|
d64b386d03 | Merge pull request 'Hand the lab design off to mesh-lab' (#5) from handoff/lab-to-mesh-lab into main | ||
|
|
a72fea5342 |
ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the boundary that decides whether the lab is worth having: a scenario that assigns overlay addresses, elects the hub and writes peer configuration certifies its own work — if the mesh's peering is broken, that scenario still comes up green. The most valuable thing the lab can test is exactly the part pre-building would replace. So a scenario declares what a hosting provider and a home router would provide: segments, which machine sits where at which address, what NAT is between them, which ports are forwarded, which machines are detached. It declares nothing about overlay addresses, hubs, peering, names or certificates, all of which become outcomes to observe. The declaration has four parts — segments, machines, place, snapshot — and the two scenario classes differ only in place. That is what makes one a strict subset of the other rather than a fork. Research 004's most important finding becomes a format constraint rather than a footnote: the routable segment must use RFC 5737 documentation space, because the mesh decides public versus private by matching the address, and a private range there makes the hub test as unreachable while the mesh silently never forms. A segment without behind: is routable, and a non-documentation address in it should be refused before anything is raised — ADR 0008 applied to a configuration file, since the failure it prevents has no error at all. Four things left open, including the one that matters most: a lab machine is always privileged, so the user and edge profiles have no scenario that exercises them. |
||
|
|
4d387d2998 |
Hand the lab design off to mesh-lab
Playbook 04: the design names its owner and flips to in-progress. code moves from hal to mesh-lab, and decisions gains 0029 and 0030 — the design now rests on three records rather than one. repos.md marks mesh-lab as the one target repository that exists. The other six remain the target, not the present, and saying so is the point: a map that lists repositories which do not exist is a map that will be believed. |
||
|
|
14f1c035a9 | Merge pull request 'The lab comes first, and its first scenario has no pipeline' (#4) from design/lab-bootstrap-scenario into main | ||
|
|
4a628bf3fe |
Issue 007: the instance is fixed, the class is the issue
The incus hook landed and does the post-install work — group, subordinate id ranges, both units, storage pool, bridge, default profile. Verified independently here: group exists with the operator in it, both id files carry the range, service active, pool reports CREATED. The only thing that fails is a shell whose process tree predates the usermod, which is how group membership works and not a defect. So the mechanism was never missing. Hooks are the right place and they work. The gap is narrower and worse: the hook did six things, six checks were then performed by a human by hand, and nothing in the pipeline asserted any of them. A pipeline that dispatched a hook which silently never fired would have been green in the same 48 seconds — and a hook named for a feature its module does not carry is skipped without complaint, thirteen of which were found at once in the past. The six manual checks are, almost word for word, the module's own verification: outcomes rather than steps, which is exactly the shape the lab design asks for. They currently live in a chat message. In the module they would run on every delivery to every node. Status moves to diagnosing rather than resolved, and fixed-by records the instance explicitly as the instance only. |
||
|
|
09489a298c |
ADR 0030: the repository structure, and the rule that names them
The tiers were settled and the product was named, but the repositories themselves existed only in a research sketch. That had already caused two problems. ADR 0029 makes the lab phase 0 of the migration and could not say where it lives, because no record named a repository. And the sketch contradicted an accepted record: it listed mesh-hq while ADR 0028 had decided novox/hq and explicitly rejected that name. A design resting on research is resting on something that can change without a decision. Corrected in the research too. The naming rule, which both earlier records implied and neither stated: a repository belonging to a product carries that product's prefix; a company-scoped one does not. That is why this repository is hq and the mesh's are mesh-*. Seven repositories recorded — host, substrate, control, surfaces, sdk, lab, and this one. The lab gets its own: it ships to nobody, outlives any single tier, and drives virtualisation on a workstation, which nothing else does. Inside the host it would couple development tooling to a shipped component; inside the control plane the bootstrap scenario would depend on a tier that does not exist when it is needed. Tier 4 is deliberately not decided. Whether the catalogue is one repository, one per domain or one per application stays open from ADR 0015 and is blocked on research 005 — how many repositories hold domains cannot be answered before knowing what the domains are. mesh-catalog appears in the sketch and is not decided by this record. The cost is stated rather than glossed: seven release cadences where there is one, and cross-repository changes that used to be one commit. |
||
|
|
b4904fec7e |
The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a complete mesh — forge, coordinator, cascade, verify. That is unusable for building the new mesh, because all four are tier 2 and do not exist yet. And research 009 had the sequence backwards. It placed the lab at phase B as verification of tiers already built, but tier 0 is the component that takes over a machine's packages, services and network. It cannot be developed against a machine anyone needs. The lab has to exist before the thing it will test. ADR 0029 splits scenarios into two classes. The bootstrap scenario is virtual machines, the host binary and a pinned bundle, with the verdict coming from what the host reports about the state it reconciled. The full scenario is the designed one. The first is a strict subset of the second — same virtualisation, same networking, same lifecycle, stopping before a control plane exists — so the second is reached by addition rather than rework. The consequence worth having: raising a node from nothing stops being the least-exercised path in the system and becomes the inner development loop. It also settles the runner's two jobs. Scenario lifecycle is needed immediately, because something must materialise and reset a mesh before anything can be written against it. Assertion execution waits for the full scenario. Corrects a stale claim in the design while amending it: it argued scenarios were affordable with system containers and would not be with virtual machines. ADR 0016 superseded that reasoning and the text had not followed. Issue 007: the lab's first requirement is installed and unusable. The virtualisation package is present and explicitly installed; both units are disabled, the operator is in no group, and the client reports the server unreachable. Not issue 001 again — that is an install failing while reporting success. This is an install succeeding when success was not the point. A package is files; a capability is a running service and an identity permitted to reach it, and the module model has no vocabulary for the second. |
||
|
|
1570234ac0 | Merge pull request 'Publish-safe: remove two disclosures, and record Nox as the answer to 006' (#3) from chore/publish-safe into main | ||
|
|
c0ae8dec96 |
Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the moment the public rule stops being aspirational. A module name identified a specific laptop model — hardware inventory, which is operational detail about one installation rather than a lesson that travels. Generalised. ADR 0028 named a forge username in a repository path, which the public rule forbids, and the sentence had also gone stale: the repository it described was subsequently verified empty of anything unique and removed. Rewritten to state what happened without the username. Removing a disclosure from a record is the same class as fixing a path — the rule that permits it outranks the one that forbids editing. Issue 006 gains its proposed direction: Nox works from within this repository rather than these documents being synced into the knowledge base. Better on three counts — no copy, so no drift; no fourth knowledge system, which was the original objection; always current. But it changes the promise, and the issue says so. ADR 0019 promised these documents would surface BESIDE everything else in a symptom search. An agent that must be asked is reachable, not surfacing, and the two differ in precisely the case the operational memory exists for — someone debugging an error with no reason to suspect HQ knows anything about it. The question narrows to whether a symptom search finds this content without the searcher already suspecting it. |
||
|
|
d0ec3e9373 | Merge pull request 'Retire the HAL name where it points forward' (#2) from chore/retire-the-hal-name into main |