fa62c7f0e4636d26d3615beaea3832231e24ac01
43
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fa62c7f0e4 |
011: correct the rule — contexts, not processes
An earlier version argued a dashboard reading a dozen stores was caught by exclusive ownership, because a dashboard is a surface and surfaces speak to an interface. Wrong, and it drew the line in the wrong place. The mesh's own board showing nodes, modules and deployments is not a separate context reaching across a boundary — it is the mesh showing its own data. Requiring it to go through an interface to reach facts its own context owns is ceremony. The rule is that a CONTEXT is granted what it exclusively owns. Everything inside it — service, surface, tools — reads that store freely. What is forbidden is a different context reading it. Which is what the consumer count already showed: the problem was never surfaces, it was three other contexts keeping their tables in the mesh's database. |
||
|
|
e71d532c2e |
011: request or subscription is derived, not chosen
Asked what the distinction actually is, and the SQL half needed correcting first: under exclusive ownership SQL runs against your own database and nothing else, whatever transport a query might travel over. Both options are the mesh's own channel and both ride the broker, so the transport is not the distinction. The distinction is where the answer lives when you need it. A request asks at the moment and waits — always current, costs a round trip, cannot answer when the other side is down. A subscription keeps a local copy — instant, works offline, as current as the last event received, and you must handle what you missed. What decides is not taste. ADR 0036 makes disconnection an ordinary situation rather than an exception, so anything that must keep working while disconnected CANNOT use a request: there is nobody to ask. And the converse — anything where a stale answer is worse than no answer cannot use a subscription. A display can lag; a decision about whether a grant is still valid cannot. So an apparently open question turns out to be derived from a decision already taken. What stays open is narrower: what a consumer does about the events it missed while disconnected — replay from a point, ask once for a full picture and resume, or rebuild. The question every projection has. Also recorded: separate databases are required in the new design, and the shared registry is a leftover rather than a pattern. |
||
|
|
7e83723b7b |
011: rewrite the question table, which had gone stale silently
Several edits to the overview matched nothing and returned success, so the question table still carried answers superseded two or three exchanges ago — "when two modules provide one name, who chooses" was still open in the table while answered in the file it pointed at, and nothing recorded the instantiation edge, instance counts, grants, bootstrap provisioning, the tool audience, or the registry consumer check. That is the fault this repository catalogues, committed by the thing cataloguing it: a string replacement that found no match, reported nothing, and left the document claiming a state it did not have. Rewritten from what the documents actually say rather than patched again. Nine questions settled, fourteen live, and the split is now visible instead of implied. |
||
|
|
afcc355744 |
011: checked the registry's real consumers, and the question was the wrong shape
The exclusive-ownership rule turned on whether every reader of the mesh registry could be served another way. Eighteen consumers open a direct connection. Four groups, and only one is work. The owner and its machinery keep reading, because they own it. The node appliers are already resolved — ADR 0037 stops the host querying the mesh database, decided for tier reasons with nothing to do with this. The bulk are FOREIGN TENANTS. The work engine holds ten of its own tables in the registry's database, the knowledge base two, pipeline logs one. Thirteen foreign tables across three contexts, which is how-we-build §4's shared schema counted. So the question was the wrong shape: the problem is not readers needing a new route to data, it is tenants needing to move out. Tasks, agents and teams have nothing to do with nodes and modules and are co-located by history. Give that context its own database and its dependency on the registry shrinks to one table. A handful of genuine cross-context reads remain, small enough to enumerate rather than estimate. The rule holds. Left open: whether those reads want an interface or events. Asking which nodes exist at the moment you need to know is a request; reacting when a node appears is a subscription, and some consumers want both. |
||
|
|
6b1aab6a1e |
011: the dashboard case, and why exclusive ownership is the tier rule
Raised as the hardest test of the rule: a board showing nodes, modules, pipelines, agents and tasks wants to read a dozen stores, and under exclusive ownership it can read none of them. It survives, and not by luck. The board is a SURFACE, and surfaces already may not do this — the skeleton puts `api/` in the control plane as the one interface every surface speaks to, and tier 3 as thin, no logic. A board reading stores directly is a surface reaching past the context that owns the data, which the tier rule forbids for reasons that have nothing to do with databases. So it is not a counter-example; it is an instance the rule catches. And the two rules turn out to be one rule seen from two sides: exclusive ownership is the tier rule expressed in terms of storage. The general shape for anything needing to see across many things: consume the record and own your own view. A reporting context builds a projection from events and reads its own store, never anybody else's. The cost said plainly rather than buried: a projection is more work than a join, and it lags. A board queries the mesh's own database directly today — ordinary, working — and this rule makes that a migration rather than a preference. The reason to pay it is §4's already-measured cost, not elegance. |
||
|
|
aa767d17a8 |
011: a module is granted only what it exclusively owns
Reconsidered by the operator — maybe shared databases should not be allowed at all — and the stricter version is better and goes further than the schemas it replaces. No shared writes, and no read-only role on another module's database either. Reading another context's tables couples you to its layout exactly as firmly as writing them does, and the coupling is harder to see because nothing breaks until the owner changes a column. That is how-we-build §4 taken at its word rather than at its letter. The permissive version — a per-consumer schema, revocable, with cross-context joins possible but deliberate — kept the letter and left the temptation. A boundary that is merely inconvenient to cross is a boundary that gets crossed. The cost is cross-module reporting, and it is the point rather than a regrettable side effect: anything wanting to know what several modules hold consumes their events or calls their interface. That is §4's whole argument, and the mesh already has both mechanisms. What gets harder is precisely the thing that was making work belonging to one context keep having to be implemented in another. And it is the first clear instance of what this effort has been hunting — what the design DELETES rather than adds. Grant kinds collapse to one: an exclusive resource. With them go the question of who owns which table, the guessing at revocation time, cross-module migration ordering, and a class of permission modelling a shared store would otherwise need. One thing it does not answer, recorded because it could make the rule unworkable: the mesh's own registry is read directly by many things today, and under this rule they consume events or call tools instead. Achievable in principle. Whether EVERY current consumer can be served that way is unchecked, and should be before this becomes a decision. |
||
|
|
4c8515507a |
011: where a binding lives follows the scope, and a grant is not always a whole resource
Two questions asked directly, and the second collides with a rule in force. One module on two nodes sharing a database corrects something stated flatly: the binding is not "recorded on the assignment". Where it is written down FOLLOWS THE SCOPE. A shared grant belongs to the module and every assignment references the same one — which is the answer for two nodes wanting one database between them. A per-instance grant belongs to the assignment. Same relation, two homes, and which home is what makes two instances share something or not. Several modules adding their own tables to one database is three needs wearing one sentence, and a provider offers KINDS of grant rather than one: a database for a consumer whose tables are nobody else's business, a read-only role for one that needs to see what another holds, and a SCHEMA within a shared database for the case actually asked about. Loose tables in a shared database is what how-we-build §4 warns against in as many words — several domains sharing one forty-five-table schema, which is why work belonging to one context keeps having to be implemented in another. Not a style objection; the observed cost, already paid. A per-consumer schema keeps what the request wants and drops what §4 objects to. Same database, same connection, same backup, and a cross-schema read remains physically possible when genuinely needed. What it adds is ownership: migrations touch one namespace, two modules cannot collide over a table name, and revoking drops the schema rather than guessing which tables belonged to whom. So the fault §4 names is still possible and no longer accidental — a cross-context join becomes something somebody deliberately writes rather than the path of least resistance. And revocation becomes answerable, which the whole-database version never was. |
||
|
|
fa9889536c |
011: tools have a different audience, migrations cross the edge, provisioning is early
Three additions, and the third kills an assumption. Tools are the most common content in the catalogue — 56 of 126 modules, more than carry a service — and they survive the split without fitting either half. A tool is not an artifact and not node state; it is a contract the mesh publishes on a module's behalf, and what consumes it is an AGENT rather than another module. That is a second audience the design has not described. Whether it is one relation with two audiences or two relations is cheap to decide now and expensive later. A migration belongs to the CONSUMER and runs on the PROVIDER. A game's migrations run against the database the store granted it: owned by the consumer, hosted inside something it does not control, ordered after the provisioning edge because there is nothing to migrate until the grant exists, and scoped to that grant. Ownership crosses the edge, which nothing in provides and requires expresses — and it gives a consumer's own install an internal order, provisioned then migrated then started, that depends on an edge rather than on its contents. And provisioning is EARLY, not late. The assumption worth killing is that it is something the control plane does for consumers once a mesh is running. The mesh's own registry database is provisioned before there is a mesh, and so is its virtual host on the broker: the store runs from the carried bundle, a database is created in it, the mesh's own schema is applied, and only then does a control plane exist. Steps two and three happen before there is a mesh to do them, so provisioning is part of the bootstrap and part of what the bundle has to express. Which strains ADR 0043. The host applies declared state ON THIS MACHINE, and a database inside a running store is not a file or a unit. At bootstrap it is at least local — the store is on the same machine. Afterwards a consumer on one node provisioned from a store on another is the ordinary case and reaching it is not the host's job. The same operation is local at bootstrap and remote later, which is either two mechanisms or one with a tier boundary crossing inside it. Currently the sharpest unresolved thing in the effort. |
||
|
|
a3c7e7e1f1 |
011: providing is a facet, and the assignment is a third thing
Any hosted service can be a factory — an identity provider grants clients, an analytics service grants a tracking identity, a mail server grants mailboxes, an application platform grants a project that is several of those at once. Providing is a FACET a module may have, not a kind of module it is, which is the same conclusion this effort reached about services and applications arriving from the other direction. So `provider` stops being a category too. Two relational stores from different vendors both grant "a database" and are the sharpest possible test of the substitutability rule. They fail it completely — different protocol, dialect, driver, client library compiled into the consumer — so `database` stays a tag, now with two real providers rather than a thought experiment. The assignment is a third entity, recorded because the operator tried the alternative: modules were once node-agnostic and it did not survive. Several of a provider's properties belong to neither end — where its state lives, how it is reached, tuning derived from the machine's hardware, which instance serves a given consumer. Not the catalogue, because they differ per node; not the node, because they are about this module. A design with only modules and nodes has nowhere to put them, which is what node-agnostic ran out of. The current system already stores environment values per module AND per node, arriving the same way. Which answers the question asked directly: two nodes both run a store, so which serves a consumer? Neither obvious answer. Not the consumer naming a node — that is placement in the consumer's manifest, a game edited because a database moved. Not the consumer not caring — for presence it genuinely does not, for instantiation it cares permanently. What the consumer knows is the SCOPE of its own need: one instance shared across every instance of itself, or one each. That decides, and needs no node named. Then the mesh binds, and the binding is recorded on the assignment and is sticky — a resolver that re-derives which store serves a consumer will one day derive a different answer and relocate a database. |
||
|
|
13c6068874 |
011: the provider shape generalises, and two things differ inside it
The broker has all nine properties the store has. So do the object store and the image registry. A substrate service is a SERVICE PLUS A FACTORY, there are four of them, and the pattern generalises past the substrate: anything granting something per consumer has this shape. Two differences matter more than the similarity. The broker cannot be managed over the broker. ADR 0001 makes it the channel every node takes work from and ADR 0039 makes it the security boundary, so the module providing it is also the way modules are managed — a declaration cannot be delivered to it over itself. Nothing else has that property; the store is consumed by the control plane but is not how the control plane REACHES anything. This is what the carried bundle exists for: the broker is raised from what the host carries because there is no other way to raise it. A constraint on one module, not a general rule, and a schema with no way to say so hides it. And two modules of identical shape want opposite instance counts. The broker is one per mesh by decision. The store cannot be, because a node that must keep working while disconnected cannot depend on a database elsewhere. Which settles what cases.md left open: how many instances is NOT derivable from what a module is. It is a per-module decision, it has to be declared, and nothing in provides, requires or excludes says it. Revocation differs in consequence too. Dropping a database leaves data until something removes it — a leak, recoverable. Dropping a virtual host loses whatever was undelivered — silent, and not. Same relation, different blast radius, which argues for the provider deciding what revocation means rather than the mesh applying one rule. File renamed: it was never really about postgres. |
||
|
|
f160b28a71 |
011: postgres worked through, and "one kind of edge" was wrong
The tidy version said a module provides names and requires names and that is the only edge. Working postgres through completely disproves it. A small game wanting to store data does not require postgres to EXIST. It requires postgres to MAKE IT A DATABASE and hand back credentials. Those are different relations in every way that matters: one creates something per consumer, carries a payload back, can be revoked, and leaves the provider holding state about who was granted what. The other creates nothing. So: two kinds of edge, one graph. Instantiation implies presence; presence does not imply instantiation. The current system already had exactly this split — `dependencies` for presence, `requires: provision:` for instantiation, with the resolver deriving one from the other. analysis.md called that derivation a convenience. It is not: it is the correct relationship between two genuinely different relations, and the design had collapsed them. Postgres also turns out to be nine things, not one. A container. Persistent state where moving nodes is a migration rather than a reschedule. Configuration partly derived from the machine's hardware. A tool surface. A provisioner. Its own bookkeeping about what it granted, which is not the data it stores. An exposure decision per node it runs on. Credentials it generates, which means a provisioning edge carries a secret. And health that is not "the container is up". Four questions the worked example makes concrete rather than abstract. WHICH postgres, when there are two — a consumer of `terminal` does not care and a consumer of a database cares permanently. How many instances a module should have, which cannot be a global rule because one-per-mesh is wrong for a store a disconnected node needs and one-per-node is wrong for the mesh's own registry. What happens to a grant when its consumer is removed, where dropping is data loss and keeping is a leak. And whether a declaration is composed PER NODE from what that node reported — because tuning follows hardware the control plane cannot know, and the alternative is the host deciding, which ADR 0037 forbids. |
||
|
|
c9c2dfe686 |
011: what a feature is, and what it splits into
The operator wants features gone, and 006 left it open. Measured, and the answer is that nothing replaces them because they were never one concept. A feature is a kind of content a module carries, detected from its directory: twenty-one of them, each with a handler owning six stages — build, publish, install, configure, start, verify. The structural finding: EVERY handler implements EVERY stage. `configs` writes files onto a node, has nothing to build, and has a build stage. `npm` publishes to a registry, has nothing to start, and has a start stage. One interface spans build-time and apply-time, so every kind of content must implement both halves and most do nothing in one — and a stage that does nothing looks exactly like a stage that failed to do anything. They split four ways, across three tiers. Artifacts built once per version and published, where no node is involved — delivery. Resources that are desired state on a machine, which is what ADR 0043 already describes and the host already does — tier 0. Actions run once against something that is not this machine, like a migration against a database on another node — delivery, and seeds go entirely. And checks: the prerequisites are REQUIREMENTS IN DISGUISE, a module saying what must be true before it can be installed, which is what an edge in the graph says; the verifiers are the read-back the host already performs. So `feature` is one word for four things spanning three tiers, which is why the pipeline is hard to reason about. One property must survive the split, and it is the thing the current design got right: content is DETECTED, relationships are DECLARED. A module that says it has migrations and has none is a fault nobody sees until it matters — but what it requires and provides is not visible in a directory and has to be said. |
||
|
|
ae099482a9 |
011: twenty cases, and two axes nothing covers
Before settling a schema, what a module can actually be. Twenty kinds of thing, with the hard ones at the end because they are the point. The ordinary nine are unsurprising: a supervised service, a system package with configuration, an application a person launches, a command-line tool, a library that never runs, a one-shot task, a scheduled one, an adapter, and a standalone application whose only difference is where its source lives. The eleven that break a naive schema are where the work is. Something that is a service AND an application — a git forge is consumed as a remote and operated through a web interface, and neither reading is wrong. Something that provides and consumes, because provider and consumer are ends of edges rather than kinds of module. Something the mesh installs that then becomes a node CAPABILITY, which means a node's provides-list is partly derived from what is installed on it and not only detected. Something that must be adopted rather than installed. Something that is a set rather than a thing. Something with exactly one instance for the whole mesh, where assigning it twice is not redundancy but two meshes. Something that is not software at all — a firewall policy, a DNS record, pure desired state, which fits the host's declaration model exactly and an installable package model not at all. An agent. The host itself, which is not a module and needs a schema that can say so. And the things the mesh depends on and does not control, which are why a node can be perfectly configured and still not work. Nine axes come out of it. Two are covered by nothing anyone has proposed: HOW MANY INSTANCES a thing may have, and WHETHER TWO CAN COEXIST — `excludes` covers part of the second and nothing covers the first. And one question the cases sharpen: is "runs" a property or a kind? The axes say property — one schema with a field saying how it runs, `never` included. The alternative is several kinds of module with different schemas, which is the taxonomy this effort already rejected once for services and applications. |
||
|
|
f55ecc1a47 |
011: an abstract name needs providers that are actually substitutable
Two corrections from the operator, and the first improves the design rather than narrowing it. `database` is not an edge. The test it fails, and the test the proposal was missing: can a consumer be switched from one provider to another WITHOUT CHANGING? A module speaking Postgres does not speak MongoDB or SQL Server — different wire protocol, dialect, driver — so a consumer declaring `requires: database` and handed any of them breaks. The name promises what no provider can deliver, and the resolver would report a requirement satisfied that is not. `terminal` passes: anything that runs a command in a terminal works and the consumer never learns which it got. So the ADAPTER is what creates an interface. `ai-assistant` is legitimate exactly because adapters normalise what is behind it. Without one there is no interface, there is a category — and a category is a TAG. Tags describe, edges bind, and keeping them apart is what stops the catalogue acquiring a second kind of relationship that looks like a dependency and is not, which is what a folder named after a domain already was. And the domain module goes. A `networking` module gathering a firewall, a resolver and a proxy under one name came from an older shape and does not fit — there is no such thing to install. There is core infrastructure: concrete modules named individually, not flavourable, with no grouping module standing in front of them. Fixed three places where the revision left the old rule standing, including an example manifest still requiring `database` — the kind of contradiction that would have been read as the design rather than as a leftover. |
||
|
|
e20a09ae80 |
011: one kind of edge
The design, rather than an account of what exists. A module provides names and requires names, and that single relation absorbs three things this effort had listed separately: requiring another module is requiring a concrete name, requiring a resource is requiring an abstract one, and an interface is simply a name with more than one provider. Nothing has to declare that it is an interface — it either has one provider or several. The move that does the most work: a NODE provides names too. Its profile is a set of them — display-server, container-runtime, an architecture — so a module requiring a display server is satisfied by the node exactly as one requiring a database is satisfied by another module. One resolution instead of two, and a graphical application cannot land on a node without a display server for the same reason, through the same code, that it cannot land without its libraries. Which makes the host's capability detection an input to resolution rather than something a person reads. It was built to be read; it turns out to be a provides-list. `excludes` is the one genuinely new relation, because it is not derivable: two modules that both provide message-bus look interchangeable when installing both would break the machine. Constraints are not placement. They say what must be true of a node, never which node — which is the mistake the measurement found in the current catalogue, where a module pins its database to a named node so a second node cannot provide it without editing the consumer. What it deletes, for the design: the module/resource distinction, the interface as a kind of thing, capability checking as a separate mechanism, domain grouping — folders assert relationships where edges record them, so a domain becomes a query over the graph rather than a directory somebody keeps true — and possibly tiers, if a tier is just a computed level. What it does not delete, stated so it is not discovered later: a resolver still has to exist, with version constraints and conflicts, and the design owes an answer on what it delegates rather than reimplements. |
||
|
|
c3a2984b3e |
011: measured, and the premise was wrong — the graph is not missing
The effort was opened to ask whether the catalogue's missing structure is a graph. It is not missing. 126 manifests, 103 edges, no cycles, nothing dangling, deepest chain of five — and a resolver in the SDK that topologically sorts them, already called by the tool loader at startup, the installer when syncing modules onto a node, and the delivery coordinator when expanding what a change affects. It already does something this effort assumed would need designing: a requirement on another module's provision is treated as an implicit edge to the module that provides it. So "ordering by the graph", which ADR 0043 makes the control plane's job, is a thing to call rather than a thing to build. The one place the graph is wrong, it is wrong about the substrate. A module needing a database declares `provider: postgres` inside `provisions:` — which is what a module OFFERS — so the resolver, which reads `dependencies:` and `requires:`, never sees it. Three edges are invisible this way, and they are the mesh's own database, the mesh's own broker, and the work engine's database. The consequence is measurable: computing what a working mesh needs from the declared graph gives registry -> sdk -> mesh -> meshware. Four modules, four levels, no database. Arithmetically correct and obviously wrong, for exactly one reason — a field that means "depends on" is not read as one. That is 04-ISSUES/003 in a new form: not a key nothing reads, but a key read as something other than what it means. Two latent defects, both contrary to ADR 0008 and both in the component ADR 0043 makes responsible for ordering a host will apply without question: a cycle warns and falls back to input order, and a dependency that does not exist warns and continues. Neither has fired, because the catalogue currently has no cycles and nothing dangling, which is why nobody has noticed. And placement is decided in the catalogue: a provision pins itself to a named node in the manifest. Which node runs what is an inventory decision — tier 2 by the skeleton's own test — so a second node cannot provide the mesh's database without editing the module that consumes it. What the graph would DELETE is currently nothing. What it would add is three declarations that no manifest uses today: excludes, a required node capability, and an interface with adapters. Whether they would be used is not measured, and zero usage is equally consistent with nobody needing them and nobody being able to express them. |
||
|
|
278f7427ed |
012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly succeeded or failed, with a severity per line. Taken with one change — the overall is DERIVED as the worst mark present, never written alongside. Two fields maintained independently drift, and a briefing reading "full success" while carrying a failed line is exactly the fault this record keeps cataloguing. An outcome computed from its lines cannot disagree with them. Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success — adoption will meet configuration it cannot parse and state it cannot read, and folding those into "fine" is the same move as reporting an installed package as a capability. And adding severity reopens something the earlier rule did not cover. "Flags inform, they do not block" was decided about CONFLICTS, where the mesh chose deliberately and the machine still works. A failure is not "we chose" but "we could not". Treating both the same makes a node where something the mesh needed never happened indistinguishable from one where a log level differed. |
||
|
|
0106318bcb |
012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning for each is the useful part. What is already on the machine stays, the conflict is flagged, adoption completes. This buys non-destructiveness by construction: the class that made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — cannot arise, because nothing tied to the machine's physical reality is overwritten. It exposes the mirror. The mesh's configuration is not only preference; some of it is what a module needs to function. Keeping the machine's version there produces a module that is installed and does not work, which is 04-ISSUES/007 arriving from a direction that issue did not anticipate. And a fleet where every node kept its own settings is one where a module works on one node and fails on another with nothing able to say why. So neither direction is right as a blanket, and the question is not whose configuration wins. It is whether the module REQUIRES the setting or merely PREFERS it — required contradictions cannot be kept without breaking the module, preferences should always yield to what is there. That is a property of the module's declaration rather than of the adoption algorithm, which makes it one more thing the graph would carry. Until modules can say which of their settings are load-bearing, adoption is defaulting in the dark, and the default chosen is the one that does not break the machine it is adopting. |
||
|
|
bcb18c7329 |
012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's disagree, the mesh's version is installed, the conflict is flagged, and it is reconciled afterwards — the mesh's configuration is known to work, the machine's is not, and a half-adopted machine is a state nobody understands. So adoption always completes and flags inform rather than block, which also settles what 'adopted with open questions' prevents: nothing. The node is a node. The original is kept, so nothing is unrecoverable. One class left open rather than folded in, because it is the one place the oldest rule in this record argues the other way. 'Known to work' is true of the mesh's configuration in isolation, not on this machine. Most disagreements are preference and overwriting them is right. A few are tied to what is physically present — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — and installing ours there does not discard a preference, it can make existing data unreadable. Restoring the configuration file afterwards does not undo that. The default is settled. The exception is not 'there is a conflict' but 'applying ours would destroy something a configuration backup cannot restore', and identifying that class is open. |
||
|
|
60736199a4 |
012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort had open with two bad answers. Nothing is taken over without keeping what was there. Adoption happens on machines somebody is already using, and the configuration being taken over is configuration somebody chose. This is a never rule rather than a courtesy, and it earns that by the same incident the mesh's strongest rule carries: the worst loss in this record came from a tool acting on a path it did not own. Adoption is that act made deliberate, which makes the safeguard obligatory. And adoption produces a briefing, not just a result. It meets things a script cannot decide — a runtime configured one way against a mesh wanting another, a package pinned for a reason, local settings the mesh has no opinion about. Silently winning is wrong in both directions and refusing outright makes a machine in use unadoptable. So conflicts are FLAGGED: what it found, what it took over, what it could not resolve, written to be read by a person or an agent as the first thing a session on that node has to work with. That is the declaration parser's principle at a larger scale — name every problem at once, to somebody who can act on it. The question it turns on is recorded rather than assumed away: are flags advisory or blocking? A briefing nobody opens is worse than a failure, because the machine is in service and the record says it went well — 04-ISSUES/003 again. Working position: the node is usable and the mesh KNOWS it has unresolved adoption questions, as a state something can ask about rather than a document in a log directory. What that state prevents is undecided. |
||
|
|
ddb8091f68 |
Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not. The host can be told to run a container or install a package; both need a file, and asking where the host gets it produced a bad trilemma — carry everything, download at apply time, or push the files in first. Downloading fails on the first node, which cannot fetch the image registry from the image registry it is trying to start. The reframing came from the operator: the machine is not offline, and what matters is WHEN the fetching happens. Move it from apply time to build time — build the installer on a machine with a network, tailored to the target, apply it on a target that then needs nothing. The same move the lab already made for its router image. Which makes the question not where artifacts come from but what is missing from THIS machine, and that needs two things answered: the closure for a one-node mesh, and how a machine already in use becomes one. Adoption is the second half, and it is sharper than it sounds. Having a package installed is not owning it: a container runtime found already present carries settings somebody chose, and noticing the binary exists discovers none of them. It was also the original path — 00-as-is/05 records adoption of a pre-existing machine's configuration as the original mechanism, since made legacy and explicitly out of scope for the lab. It returns for a different reason than it was dropped for. Two collisions recorded rather than discovered later. ADR 0004 has managed files generated and never edited, and adoption needs a one-time import before that rule starts applying — three states, and the middle one is new. And ADR 0043 says the host never touches what it did not create, which is exactly what adoption does; that rule needs a companion rather than an exception. Eight open questions, including whether 'tier' is just a coarse view of a graph level, whether owning a package means owning its version, and what cannot be precomputed at all — because tailoring moves the cost of building from source rather than removing it. |
||
|
|
b9facf9375 |
Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply declared state on this machine — and the six absorbed concerns are instances of it, not additions to it. Specifies the six parts and what each owns, and the two properties that make apply trustworthy rather than merely present: every applier reads back, because setting a value is not evidence the value took; and what was applied is recorded after it works, never before, because a failed apply leaves the machine wherever it reached and nothing must claim otherwise. Build order is staged so each stage is verifiable in the lab before the next exists. Stage 1 is profile and inventory — no control plane, no declarations, no network — and it is deliberately the smallest useful thing, because `place:` has nothing to place and the lab therefore raises empty machines. Stage 1 ends that, and every later stage is tested by a lab that already works. Stage 2 is the one that could invalidate the tier boundary: whether one host can raise the substrate alone is Move 1's assumption and has never been proved. Every decision the design rests on is given the test that asserts it, per 0034 — including the dependency-direction lint, which is what makes "the host never queries the mesh database" a rule rather than an intention. Six things left open and named, including the one that host-size.md could not measure: zero dependencies, but still six vocabularies. |
||
|
|
902739acb6 |
Research 011 — the module graph
The proposal to split modules into provisioning services and applications was worked through and abandoned, for a reason worth keeping: it cannot be filed consistently. A git forge is consumed as a service and operated through a web interface; an analytics service grants tracking identity and is a dashboard. The operator's correction is the sharper form — what runs on the machine is a supervised container, not something a user started. That is a fact about HOW a thing runs, not about what kind of thing it is. So it is a facet, and 0002 survives: everything is a module. What the catalogue is missing is not a taxonomy but a graph. Grouping asserts relationships; a graph records them. Five declarations, of which two exist: requires/provides a resource (yes), requires/excludes another module (no), requires a node capability (no). Plus interface modules that carry no implementation, with adapters providing them. Recorded because it matters: this is a package manager's model, and pacman already has all of it — depends, conflicts, and provides as virtual packages, which is exactly the interface/adapter idea. Arriving there independently is evidence for the shape. It is also a warning about what not to reimplement. Working position on capabilities, to be tested: intrinsic ones (hardware, architecture, network position) are detected and never installed, and a module requiring one it lacks is impossible rather than unresolved. Provided ones (a display server, a container runtime) are not a separate kind of thing — they are modules that provide a capability, so "may the mesh install a capability" is not policy, it is dependency resolution. Issue 007 then bears directly: an installed package is not a capability. Also captured: the operator's assessment that the machinery around a module — scheduled tasks, hooks, migrations, config and env — is worth keeping, seeds are not, and the integration is wrong enough to need a major refactor. Research 005 found supporting evidence from another direction, that the densest apparent coupling in the catalogue is manifest boilerplate churn. The first open question is the one that decides whether this is progress: what does the graph DELETE? If modules gain declarations and lose nothing, it is motion. |
||
|
|
72b22830f3 |
ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation. The open question posed a class distinction — full nodes and lesser presences. There is none. Reachability is state, not kind, which promotes the host's local store from a component to a requirement: it is what makes disconnection ordinary rather than exceptional. The reduced contract the question reached for is real but it is capability, and that belongs in the profile. 0037 (accepted): the host applies, it does not decide. Measured rather than argued — the absorption is smaller than the machinery that already applies state, and eight of ten adapters carry no dependency to move. The two that do open a Postgres connection to the control plane, which inside tier 0 is the one thing the tier rule exists to forbid. So each concern splits: deciding needs every other node and stays in tier 2; applying needs root and locality and goes to tier 0. The host carries ONE concern, of which the six are instances. 0038 (proposed): a node joins by linking first. The operator's two-modes proposal, adopted as intent and corrected as structure. Two modes is two code paths where the first runs once per mesh and rots — and the mesh already has that fault in its worst form, as three hand-run shell scripts. Instead: one behaviour, two sources of declaration. The first node is not a different kind of node, it is a node whose mesh is not up yet, and its specialness is temporary and self-erasing. 0038 also shrinks the migration 0037 called expensive: a joining node never needs mesh-wide state, because the hard part of the overlay is only needed to compute the WHOLE mesh. It needs one peer. The rest arrives. Left open and said so: what may be pushed over the link and how a joining node proves it is entitled to join, and whether one host can raise the substrate alone. |
||
|
|
42bce02bba |
006: answer the host-size question by measuring it
The skeleton's biggest unproven claim was that absorbing six concerns makes a binary whose whole argument is having no dependencies carry six of them. Measured against origin/main, and the question turns out to ask about the wrong axis. By size the absorption is SMALLER than the machinery that already applies state on a node — 2755 lines of adapters against 3059 lines of meshware, env-sync and config-sync. The host is not a new large thing; it already exists, spread across three core modules. The real risk is direction, and it is two modules wide rather than six concerns wide. Eight of ten adapters already receive derived state and only apply it, so absorbing them moves code that has no dependency to move. Two — wireguard and traefik — open a Postgres connection to the control plane and compute their own configuration, which inside tier 0 would be an upward dependency and is exactly what the tier rule forbids. And the split has already been happening without being named: dnsmasq-app needs the same node data as wireguard and does not query for it, because hand- duplicated state went wrong and someone derived it centrally instead. Eight of ten adapters are on the far side of that migration. So the absorption is not a move, it is a split: deciding stays in tier 2, applying goes to tier 0. The claim survives with its scope corrected — the host carries ONE concern, apply declared state on this machine, of which the six are instances. Stated open rather than glossed: the two unsplit modules are the two hardest, six concerns is still six vocabularies even at zero dependencies, and what the host must carry versus find is issue 007 and unresolved. Question B also recorded as answered by the operator — a node is a managed machine, and a disconnected node is still a node in a different situation. The question posed a class distinction; there is none, and what varies is state. |
||
|
|
4bf7a35568 |
Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design, design before build — and skipped for the bookkeeping. This closes that. 004 graduates. Its one open item was "not yet stood up"; the lab is stood up, and the substitution the effort turned on is now enforced by the validator before anything is raised rather than left as a thing to remember. Its certificate conclusion has a home in 01-end-to-end-testing and is designed but not built — implementation is a third axis, and an effort graduates on its conclusions. One item leaves 004 without a home and is recorded rather than lost: the reverse proxy does not set caServer, so it defaults to the production endpoint. The two lab designs read `designed` while running in production of a sort, so they become `in-progress`. And the lab gets an as-is document, which it did not have. It records what runs including the parts nobody would choose again: that `place:` is refused and the lab therefore raises EMPTY MACHINES, that the drawing shipped with no design document behind it, that a router is tagged as a machine for a reason found by a bug, and that the integration suite raises two of five scenarios while both faults found so far lived in the three it does not. 006 stays active, deliberately. Two of its open questions ARE the tier 0 design — whether absorbing six concerns makes the host too large, and whether an unprivileged node earns a place in the inventory. Playbook 04 is explicit that an open question is a reason to research, not to build around. |
||
|
|
e88b448145 |
The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing 1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle falls from ~90s, unbounded at worst, to ~15s dominated by a boot that cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write and did not without it. The consistency matters as much as the speed: three consecutive snapshots took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never finished. One honest counter-observation recorded: launching onto the fresh copy-on-write pool was slower, 20.2s against 14.3s, because the image had to be unpacked into a pool that had never seen it. Paid once per pool, and dwarfed by what snapshotting saves, but it went the other way. Doing the measurement produced the answer to how the lab installs on a clean machine, because both failure modes appeared while doing it. Installed is not available: the daemon was present with units disabled and no group. Issue 007. Available is not adequate, and this is worse: with the storage tooling absent everything worked and snapshots were seventy-six times slower. Nothing failed, nothing warned. That is a variant the mesh has not catalogued — its usual failure is reported success and did nothing; this is reported success and did it seventy-six times slower, which no error surface catches because nothing is wrong. So the lab verifies CAPABILITY, never installation, and refuses to run degraded rather than warning — a warning about a slow inner loop is read once and ignored forever. Prerequisites may arrive from a mesh module or from the lab's own bootstrap, and the second path is required rather than convenient: a lab installable only by a mesh cannot host the development of the mesh that installs it. The lab is the second thing installed by hand, after the node host, and for the same reason: something has to be first, and pretending otherwise produces a circularity papered over by a script nobody exercises. |
||
|
|
98bcd5cc49 |
Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than arguable, so it was measured. One virtual machine on a workstation with hardware virtualisation and NVMe. Raising: the launch call returns in 3.4s, the machine is actually usable after 14.3s. The gap is a design constraint — raise must wait for the second number, because reporting the first would be transport reported as effect, which is the mesh's own recurring failure. Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full copy; nothing is shared. Restore: 10.4s, usable again after 20.1s. The second snapshot exceeded two minutes and never completed. That is the more troubling number: snapshot cost here is not merely high, it is unpredictable, and a loop with a variable multi-minute step is one nobody trusts. Projected to a four-machine scenario, a reset-and-rerun cycle is about a minute and a half at best and unbounded at worst, before any of the mesh's own work begins. That is too slow for an inner loop, and ADR 0029's whole argument — that making the bootstrap path the inner loop turns the least-exercised code into the most-exercised — holds only while resetting is cheap. The cause is not virtual machines. Hardware virtualisation is present and machines boot in fourteen seconds. It is that the daemon offers exactly one storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The btrfs kernel module is available; btrfs-progs is simply not installed, which is the entire reason the driver is absent. The copy-on-write comparison was deliberately NOT run, because running it would mean installing a package by hand — which the rules forbid and which would have made the measurement unreproducible. So the honest statement is that the current configuration is too slow and the likely fix is known but unverified, rather than that btrfs fixes it. |
||
|
|
a72fea5342 |
ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the boundary that decides whether the lab is worth having: a scenario that assigns overlay addresses, elects the hub and writes peer configuration certifies its own work — if the mesh's peering is broken, that scenario still comes up green. The most valuable thing the lab can test is exactly the part pre-building would replace. So a scenario declares what a hosting provider and a home router would provide: segments, which machine sits where at which address, what NAT is between them, which ports are forwarded, which machines are detached. It declares nothing about overlay addresses, hubs, peering, names or certificates, all of which become outcomes to observe. The declaration has four parts — segments, machines, place, snapshot — and the two scenario classes differ only in place. That is what makes one a strict subset of the other rather than a fork. Research 004's most important finding becomes a format constraint rather than a footnote: the routable segment must use RFC 5737 documentation space, because the mesh decides public versus private by matching the address, and a private range there makes the hub test as unreachable while the mesh silently never forms. A segment without behind: is routable, and a non-documentation address in it should be refused before anything is raised — ADR 0008 applied to a configuration file, since the failure it prevents has no error at all. Four things left open, including the one that matters most: a lab machine is always privileged, so the user and edge profiles have no scenario that exercises them. |
||
|
|
09489a298c |
ADR 0030: the repository structure, and the rule that names them
The tiers were settled and the product was named, but the repositories themselves existed only in a research sketch. That had already caused two problems. ADR 0029 makes the lab phase 0 of the migration and could not say where it lives, because no record named a repository. And the sketch contradicted an accepted record: it listed mesh-hq while ADR 0028 had decided novox/hq and explicitly rejected that name. A design resting on research is resting on something that can change without a decision. Corrected in the research too. The naming rule, which both earlier records implied and neither stated: a repository belonging to a product carries that product's prefix; a company-scoped one does not. That is why this repository is hq and the mesh's are mesh-*. Seven repositories recorded — host, substrate, control, surfaces, sdk, lab, and this one. The lab gets its own: it ships to nobody, outlives any single tier, and drives virtualisation on a workstation, which nothing else does. Inside the host it would couple development tooling to a shipped component; inside the control plane the bootstrap scenario would depend on a tier that does not exist when it is needed. Tier 4 is deliberately not decided. Whether the catalogue is one repository, one per domain or one per application stays open from ADR 0015 and is blocked on research 005 — how many repositories hold domains cannot be answered before knowing what the domains are. mesh-catalog appears in the sketch and is not decided by this record. The cost is stated rather than glossed: seven release cadences where there is one, and cross-repository changes that used to be one commit. |
||
|
|
b4904fec7e |
The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a complete mesh — forge, coordinator, cascade, verify. That is unusable for building the new mesh, because all four are tier 2 and do not exist yet. And research 009 had the sequence backwards. It placed the lab at phase B as verification of tiers already built, but tier 0 is the component that takes over a machine's packages, services and network. It cannot be developed against a machine anyone needs. The lab has to exist before the thing it will test. ADR 0029 splits scenarios into two classes. The bootstrap scenario is virtual machines, the host binary and a pinned bundle, with the verdict coming from what the host reports about the state it reconciled. The full scenario is the designed one. The first is a strict subset of the second — same virtualisation, same networking, same lifecycle, stopping before a control plane exists — so the second is reached by addition rather than rework. The consequence worth having: raising a node from nothing stops being the least-exercised path in the system and becomes the inner development loop. It also settles the runner's two jobs. Scenario lifecycle is needed immediately, because something must materialise and reset a mesh before anything can be written against it. Assertion execution waits for the full scenario. Corrects a stale claim in the design while amending it: it argued scenarios were affordable with system containers and would not be with virtual machines. ADR 0016 superseded that reasoning and the text had not followed. Issue 007: the lab's first requirement is installed and unusable. The virtualisation package is present and explicitly installed; both units are disabled, the operator is in no group, and the client reports the server unreachable. Not issue 001 again — that is an install failing while reporting success. This is an install succeeding when success was not the point. A package is files; a capability is a running service and an identity permitted to reach it, and the module model has no vocabulary for the second. |
||
|
|
c0ae8dec96 |
Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the moment the public rule stops being aspirational. A module name identified a specific laptop model — hardware inventory, which is operational detail about one installation rather than a lesson that travels. Generalised. ADR 0028 named a forge username in a repository path, which the public rule forbids, and the sentence had also gone stale: the repository it described was subsequently verified empty of anything unique and removed. Rewritten to state what happened without the username. Removing a disclosure from a record is the same class as fixing a path — the rule that permits it outranks the one that forbids editing. Issue 006 gains its proposed direction: Nox works from within this repository rather than these documents being synced into the knowledge base. Better on three counts — no copy, so no drift; no fourth knowledge system, which was the original objection; always current. But it changes the promise, and the issue says so. ADR 0019 promised these documents would surface BESIDE everything else in a symptom search. An agent that must be asked is reachable, not surfacing, and the two differ in precisely the case the operational memory exists for — someone debugging an error with no reason to suspect HQ knows anything about it. The question narrows to whether a symptom search finds this content without the searcher already suspecting it. |
||
|
|
87f4f29cc6 |
Novox Mesh, Nox, and HQ becomes company-scoped
ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was never chosen: it arrived with the dotfiles repository this grew out of, it is borrowed, and it is borrowed from the canonical untrustworthy machine intelligence, which is an odd flag for infrastructure trusted with credentials. Timing is the substance of the decision, not an aside — the skeleton is not built, so renaming costs a search and replace now and a migration later. Nox is an identity of Novox, and specifically the agent of the MESH rather than of a node. Nodes keep their own identities. Nox addresses them, and a human mostly talks to Nox — which makes it the concrete form of the mission's vision: state an intent, and the mesh works out which node holds the thing. It holds no private channel. The gap this opens is recorded: ADR 0012 binds every agent to a home node, and a mesh-scoped agent has none, so the model needs extending. ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first product. Checked rather than assumed: the company organisation already holds live projects that the mesh builds and deploys, so they are tenants rather than peers, and the mesh is the ground they stand on. There is also company work outside the mesh already, which strengthens the case and means the eventual split is closer than "some day" — so each document's scope is fixed now, in a table, making that split mechanical instead of archaeological. The folders are deliberately not restructured yet. The skeleton takes the new vocabulary: mesh-host, mesh-substrate, mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services now that identity is a hosted workload rather than a dependency. Research 009 opens the migration, with the reframing that lowers its risk: replace the control plane, do not move the workloads. Their data never moves, so it is re-declared rather than adopted — which keeps adoption out of scope, as the lab design requires. Self-hosting is the last phase, or a failed cutover takes away the means to fix it. |
||
|
|
daf3e17c32 |
self-hosting, provisioning and delivery efforts, and the dotfiles origin
The identity provider is settled as not-substrate: the mesh does not require one, tier 2 authenticates natively, and it is a hosted service like any other. Four substrate services, not five. The tier test's second step gains the verb that matters — can the control plane START without it, not function fully without it. That verb answers the forge and the registries. They are not substrate and they are not duplicated: the control plane starts and manages nodes without a forge, it just cannot change itself. One gitea module, tier 4, and the mesh's own instance is distinguished by what it is bound to rather than by being a different module — the same answer as postgres, from the same test. It also buys a property worth having: if the forge dies the mesh keeps running. Delivery needing them is not an upward dependency, resolved the way the constitution already says to: tier 2 declares requirements, tier 4 provides implementations, the binding is data. The mechanism is provisioning, and the new idea is that the control plane is itself a consumer. Self-hosting therefore becomes a state the mesh REACHES, not a precondition. A first node comes up from pinned external artifacts and re-binds to internal providers once they exist. Today's mesh assumes the second state from the first moment, which is why the first-node path needs a script that papers over an impossibility and is the least-exercised code in the system. Made explicit, the transition is also reversible. Research 007 and 008 opened for the two areas flagged as important and complex, scoped from the weaknesses the as-is layer already documents rather than started blank. And the origin: this began as a dotfiles repository. The first two days adopt dotfiles, add per-node overrides, and introduce service symlinking with an ignore file. The flat one-directory-per-tool catalogue, linking over copying, adoption of already-configured machines, per-node overrides and the desktop modules are all inherited rather than chosen for a mesh. That is the single most useful fact for anyone changing the catalogue, it strengthens ADR 0018 — the case for links was never made for a mesh — and it explains research 005's silent fifty: dotfiles-era entries for one tool never shared a domain because they never had one. |
||
|
|
7a20358113 |
research 006: the code skeleton, and where postgres lands
A tier test as a decision procedure — five ordered questions, first match wins — so placement is answerable rather than argued. Postgres was the test case and the naive answer is wrong. Not twice, once: the control plane cannot exist without a relational store, so it is tier 1 and lives in hal-substrate/store/postgres. What differs between the mesh's own database and a project's is not the module but how that instance is brought up — pinned bundle applied by the host, versus the ordinary delivery and provisioning path. Tier is a property of the module; the bundle is a property of the mesh's own instance. The naive answer would also have made substrate reach up into the catalogue, which the dependency rule forbids. Working the test across the catalogue surfaces a third fate that neither of research 005's options covers, and it is the most common one: absorbed into the host, ceasing to be a module at all. That explains 005's one positive measurement rather than confirming it — the reachability cluster is not four modules that should be one domain module, it is four facets of one thing the host should own, expressed as modules because a module was the only unit available. Under this skeleton the overlay and firewall modules stop existing. It also partly answers the silent fifty: several are host concerns, so silence was the right signal and grouping was the wrong inference. Flags rather than settles: the identity provider is a genuine boundary case (four substrate services or five), and absorbing six concerns into a binary whose argument is that it has no dependencies is the skeleton's biggest unproven claim. |
||
|
|
00b8398d07 |
skeleton: hal-agent -> hal-host, and the two gaps the questions found
Agent is a first-class concept here — a participant, some of whom are human, holding identity and memory. Using it for the tier-0 node binary put both meanings in one document: hal-agent at tier 0, agents at tier 2. That is the anatomy-naming failure again, an evocative domain word pointing at infrastructure, and how-we-build 4 exists to catch it. ADR 0015's own title supplies the fix: the mesh brokers, nodes host, agents think. So the control plane brokers (hal-mesh), the tier-0 binary hosts (hal-host), the participant thinks (agents, untouched). hal-node was rejected — Node is the inventory aggregate, the binary is what runs on it. Recorded in the document as a near-miss rather than quietly corrected. Tiers now explained before the tree, as a boot narrative, with the point they were carrying made explicit: dependencies point only downward, that is the whole bootstrap answer, and the current mesh violates it — the database is a module, modules come from the pipeline, the pipeline needs the database. Connectivity gains the part that was missing. Naming a context says who decides, not who runs it: overlay membership is tier 0 in the host, policy is tier 2, machinery is tier 4 modules. And the host's link to the control plane deliberately does not run over the overlay, or the overlay would have to exist before a node could be told how to join it. 'tier' is this document's coinage and did not land on first reading; that is now an open question rather than settled vocabulary. |
||
|
|
b4365d8aa4 |
research 006: the mesh designed from nothing
A skeleton laid out against the stated requirements rather than derived from the current shape: four tiers, repositories at the root, and a dependency rule that only points downward. Four moves the current shape does not have. The substrate is applied by the agent from a pinned bundle, not delivered by the pipeline. That is the bootstrap circularity removed rather than worked around — the first node is the ordinary path with no control plane on the other end, which also makes it the cheapest lab scenario instead of the one nobody exercises. One agent binary with a detected capability profile — managed, user, edge. A phone becomes a capability question rather than a platform question, so it needs no second implementation. Modules declare which profiles they can land on, and an impossible assignment fails at declaration. Connectivity becomes a context. ADR 0015 names nine and none owns the overlay, resolver, firewall or ingress, while research 005 measured reachability as the only cluster in the catalogue that genuinely changes together under one intent. Gap and evidence point the same way. That is an addition to an accepted record, so it needs its own record and is not written here. Feature splits into artifact (built once per version) and part (selected per node). The conflation of those two cardinalities under one word is what makes the delivery pipeline hard to reason about. Also makes explicit in how-we-build that the main-branch rule covers this repository too. The rule already said 'without exception'; nothing was amended, so nothing is recorded. |
||
|
|
143f8ab2f1 |
research 005: which modules actually change together
The domain-grouping premise is testable, so it was tested before drawing a list. Co-change across the full history of the module catalogue, current modules only, platform namespace excluded. Nine commits in ten touch exactly one module, and 50 of 89 modules have never been edited alongside anything. Two clusters exist above that floor. Reachability holds up: proxy, resolver, firewall and VPN genuinely move together under one intent, three times in recent history. That is the shape ADR 0017 describes and the only place the measurement finds it. The provider cluster does not, and this is the finding worth having. Every multi-provider commit is a cross-cutting manifest change applied N times — feature detection, hook conventions, volume binds, network scoping. None is a change to what a database is. Merging them would not have prevented one of those commits, and the history already shows the fix that worked: verify by shape in the SDK rather than copying a script into every module. Move the concern into the machinery, do not merge the modules carrying it. ADR 0017 keeps its principle and gains a pointer to this narrowing. |
||
|
|
3f6d939930 |
Every decision is a record; the ledger is gone
papa-hq has no ledger. Its root is AGENTS.md, CLAUDE.md, README.md, every decision is a numbered record, and its graduation playbook has no path for an unrecorded decision. hal-hq now matches. The ledger's 41 entries classified as: 10 restating a record, 11 restating design docs, 15 describing how this repository works with the reasoning sitting in a README rather than anywhere citable, 3 small rules with no home, 2 superseded stubs. Mostly a copy — and a hand-maintained index, the exact pattern ADR 0022 had just rejected for the decision index on the grounds it drifted after one addition. Keeping one copy of that while removing another is not a position. It also collided by name with 02-DECISIONS/ in any directory listing. Nothing was dropped. Records 0019-0025 give the repository decisions the reasoning they never had: HQ is its own repository and is public, design has two layers, work moves through playbooks, status lives in frontmatter, issues have a front door, the numbering is the flow, HQ is the source of the constitution. 0026 records the ledger's own removal. The three orphan rules went to how-we-build, where a rule is enforced and keeps the incident that earned it — the package rule was genuinely unwritten anywhere. Two lab decisions stated only in the ledger went into the lab design. "Deliberately not decided" went to the research effort and design document each question actually belongs to. The chronological view the ledger provided is now generated from record frontmatter, which is what it was for. The cost, stated in 0026 rather than glossed: a record is more work than a table row, so the risk is a small decision going unrecorded because nobody wanted to write a document. how-we-build takes rules cheaply, which is the mitigation, not a solution. |
||
|
|
c0b35652d0 |
The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist. |
||
|
|
f05e4a0dce |
Follow papa-hq's research convention; the mesh links nothing
Research efforts move from status.md to 00-overview.md with active / graduated / abandoned, matching papa-hq so the two repositories read the same way. Playbooks, skills, README and the ledger follow. Reverses yesterday's withdrawal of the symlink note in GENESIS. The note was right and the withdrawal was wrong: the intent is that the mesh creates no symlinks at all, so a founding document listing "symlinks, not copies" as a design principle does point the opposite way from where this is going, and that is a contradiction rather than a stale detail. ADR 0018 records the position, proposed. ADR 0011 stays as it is — it is the historical decision and the incident behind it is why anyone believes either record — and is superseded in intent, not edited. Its one editorial line, which called the wider reading false, is corrected to state what is actually true: centralising who may link narrowed the incident class without closing it, because a link the installer makes resolves exactly like one made by hand. The argument that kept linking was staleness. ADR 0004 removed it: every managed file is already derived and reconciled, so a copy is the natural form and a pointer into source is the shape the mesh's own model forbids everywhere else. What is not settled, and is marked open, is how staleness gets detected — which is the decision that makes or breaks it. |
||
|
|
702efca6bb |
Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the decisions were about, and an as-is claim had nowhere to live except inside an intention. Adds 02-DESIGN/00-as-is — eleven documents written from the implementation and the operational record, not from intent, including the parts nobody would choose again. The two existing designs move under 01-to-be. Layers are declared in frontmatter and never mix: a design that ships does not move, its as-is counterpart is written, and both stand. Back-fills adr/0001-0014 for decisions taken in implementation and never recorded — the broker, the module abstraction, the mesh database, managed files, provisioning, migrations, the workspace removal, failing loudly, the constitution, application placement, linking, the employee model, the artifact, the three silos. Each marked reconstructed, dated from the history, and citing the evidence it was recovered from. The two existing records renumber to 0015 and 0016 so the ledger runs oldest first; 0017 extends 0015 to modules outside the core, principle only — the domain list is deliberately not invented here. how-we-build.md becomes the source of the mesh constitution, with a sync playbook, so the enforced copy stops being the only one that is true. Process becomes explicit: five playbooks, eight thin skills that defer to them, a repository map, and AGENTS.md with CLAUDE.md as its include. The five Observations become 04-ISSUES 001-005 where they can be owned and closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge base. That claim is what decision 27 rests on, it was never checked, and the README now says so instead of repeating it. Also corrects the ADR index into something generated, the "02-DESIGN is empty" claim, the VISION.md pointer that did not survive the repo split, and a note asserting the symlink rule was contradicted — it was a misreading; the rule forbids hand-made links, the installer links by design. |
||
|
|
cf9357e8e9 |
HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log. |