Commit Graph
26 Commits
Author SHA1 Message Date
jschoubben c171c64d3a Genesis lets the control plane state its own facts
A mesh raised from nothing must be able to say what it applied and what a machine refused (novox/hq
ADR 0134), and the first user list is what permits it. Kept in step with the controller's own
composition by the test that reads this file.
2026-09-28 16:07:20 +02:00
jschoubben 45b9a507a1 Genesis lets the controller ask any module's tool
The controller is the way in for tool calls (novox/hq ADR 0095); the first user list must say so,
or a mesh raised from nothing refuses its own first ask. Kept in step with the controller's
composition by the test that reads this file.
2026-09-28 04:21:22 +02:00
jschoubben e212da6bd2 Genesis: the first user list lets the controller hear the forge's merges 2026-09-28 03:13:16 +02:00
jschoubben 15f0dabf32 Genesis: the first user list lets the controller hear its consumers, and the account has JetStream
Two things the controller's own composition now derives and the installer's
carried list did not: the delivery subjects of the controller's consumers, and
JetStream enabled on the account — without which the first bound consumer is
refused. Found live on a mesh moved rather than raised; a fresh genesis would
have met both at first start.
2026-09-28 01:46:14 +02:00
jschoubben 6a3435629e A foundation template that raises the mesh on the bus being built
The same twelve steps, with the difference that matters: the mesh composes its own user
list and at genesis there is none, so this carries the first one — the controller's
account at a bootstrap password, rotated with the store's and replaced by the
controller's own composition from its first start.

The server's settings and the user list are separate files in one directory. Separate
because the settings belong to whoever raises the server and the users belong to the
mesh; in one directory of necessity, because an include path resolves relative to the
including file's own directory, so an absolute one sends the server looking underneath
that directory and it refuses to start.

No `verify` in the TLS block. That makes the server demand a client certificate and
nothing in the mesh presents one — a host pins this server's exact certificate and
authenticates with a password.

The controller's permissions here are checked against what the controller derives, by a
test in its own repository reading this file. They are two statements of one fact, and a
template that granted less than the controller needs would produce a mesh that comes up
and is refused on its first act.
2026-09-27 16:39:32 +02:00
jschoubben c32eada62b The base filter opens the bus and the registry in the input chain too
A container on the machine dialling a port the machine publishes reaches it
through the runtime's proxy — input, not forward — and the builder could not
reach the broker. The derived ruleset opens the mesh's own ports in both
chains; the base one now does the same.
2026-09-21 12:28:57 +02:00
jschoubben b72b71a989 The host applies the newest declaration, a file may be created once, the foundation filters first
031: a window of unacknowledged declarations is drained to the newest; the
rest are set aside and reported as superseded. 035: a file resource may say
create-once — written when absent, kept untouched when present (ADR 0087).
054: the bundle installs nftables and loads a base ruleset before the store
and broker, in the table the filter module later replaces (ADR 0088).
2026-09-21 12:11:52 +02:00
jschoubben 56124c38b7 Phase 3.2: the foundation broker binds amqp (5672) mesh-wide for consumers
Like the store, the amqp provision is plaintext 5672 (vhost-per-login), so a
consumer must reach it — bind 0.0.0.0 (firewall-gated to `mesh`, WireGuard-
encrypted on the wire) instead of loopback. amqps (5671) was already mesh-wide;
management (15672) stays loopback for the host-networked provisioner.

Issue 051 (WBS 3.2).

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:24:13 +02:00
jschoubben 9109a8c178 Phase 3.1: adopt the foundation store as the postgres module
InstallStore turns the mesh-store the foundation raised at genesis into the
postgres module, adopted in place: it verifies the module's server names the
same container and the same image the foundation is running (fail-fast on a
drift, rather than tearing down the mesh's store), then registers, builds the
provisioner, and carries the superuser in via secret accept — the mesh cannot
invent a credential that already made the databases (mirroring the control
plane's store-connection delivery, control.go). pinImage generalised to any
module for reuse.

Issue 051 (WBS 3.1). One server holds the controller's contexts and every
module's database.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:04:06 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 8e12b3c9e4 Name the decisions these tests defend, and check the bundle at all
From auditing the decision records: of 28, only 12 were named by any
test, so "which decisions are defended" could not be answered without
reading everything. ADR 0017 says a test names the decision it defends —
that rule was itself unenforced.

Most of the gap was citation, not coverage. Drift detection was tested
in several places without naming ADR 0011; the archive refusal without
naming 0012; forged declarations without naming 0002. Named now, so the
question is answerable by grep.

The bundle was the real gap: nothing tested substrate-first-node.lock at
all. It is what a machine becomes when there is no mesh to ask — the one
declaration applied with nothing to verify it against — and it was
edited by hand and read by nothing but a running host.

Two tests now assert what it carries: exactly postgres, lavinmq and the
control plane. That defends ADR 0028, which removed the object store
from the substrate after it had been a member for months on the strength
of "it cannot grant itself a bucket" — true, and the answer to only half
the test. Nothing counted what the bundle held.

Fault-injected, and the first attempt did not bite: the injection landed
on a comment line, which stripComments discards. Injecting into the
image field fails as it should.
2026-08-31 17:33:34 +02:00
jschoubben 5bc0006e83 The substrate raises a third context's store
Each context owns its own database (novox/hq ADR 0008), so a third context is a
third database, created and named the same way — which is the whole of adding
one to the bootstrap, and is why the count is not something the substrate has
an opinion about.

The schema step verifies all three now. It checked two while creating three,
which would have reported success for a context whose tables were never made.
2026-08-31 02:52:26 +02:00
jschoubben efa2e13513 The store's readiness is checked over TCP, not the socket
While the store initialises it runs a temporary server on the unix socket only,
then stops it and starts the real one. A socket check sees that temporary
server, the action exits happy, and the verify a moment later lands in the gap
between the two and fails — reported as "the action ran without error and its
own verify still fails", which is true and names nothing.

Intermittent, so it read as a slow machine. Both the action's own loop and its
verify now ask the same question, over the port the init phase deliberately
does not open: an action and its verify asking different questions is an action
that can succeed into a state its verify rejects.
2026-08-31 01:01:00 +02:00
jschoubben b2ecf25594 When the store does not come up, say what it said
The bootstrap's readiness wait failed on a loaded machine and reported
`docker exited 1:` with nothing after the colon. The file's own comment
already records this failing three times before and being fixed by
running it again — "the worst kind, because it teaches people to run
things twice".

Raising the timeout a second time would treat the symptom. What makes a
retry the only available response is a timeout that reports nothing, so
the wait now prints what pg_isready says and the store's own last lines
before giving up.
2026-08-30 04:06:10 +02:00
jschoubben bdc9c436b4 The token says what the mesh calls this machine
Found by raising a mesh end to end for the first time. Enrolment's own
help says the token "is the only thing it needs", and it also needed
--name, with no default. Without it the failure is:

  cannot reach the broker at 192.0.2.10:5671 as : username or password
  not allowed

An empty username, and nothing about the cause.

The node cannot work its own name out. The broker account it
authenticates as is named after it and exists before this machine has
been told anything, so the name has to arrive with the rest. It is not a
secret and the issuer already knows it.

--name stays, as an override for a token issued before the name
travelled in one, and says so when it is needed rather than failing at
the broker.

Also corrects the bundle example, which claimed to stop before the
control plane runs and has raised one for some time. A comment about what
something does not do is a comment nobody updates.
2026-08-30 02:36:42 +02:00
jschoubben ef400b8c66 Point the example bundle at the current control plane 2026-08-29 22:16:12 +02:00
jschoubben 06f393f9aa Point the example bundle at the current control plane 2026-08-29 22:00:06 +02:00
jschoubben 4bff67ec69 A machine waiting to be enrolled is not a broken one
The launcher already ran `host run`, and `run` on a machine with no identity
exited with an error. So a freshly installed host, sitting exactly as intended
waiting for somebody to bring it a token, would have counted three failed
starts and rolled back its own installation.

It waits now, and says what it is waiting for. That is the *hosted* state from
the lifecycle: the host is running, it has no identity, and there is nobody to
link to. Every machine passes through it.

An identity that exists and cannot be read is still a fault rather than a wait.
Treating that as "not enrolled yet" would leave a node sitting quietly for ever
while the mesh believes it is a member.

Also: a node now says it is there once a minute. Nothing but its name, because
anything more would be a report, and reports are rare where this is constant --
reading one as the other would make a quiet node look like a stale one. Not
published mandatory, unlike a report: losing one is nothing, the next is a
minute away, and the mesh reads a gap rather than counting arrivals.

Verified in the lab: a node was stopped and the mesh said "out of touch 4m",
then it was started and the mesh said "here" again, without anything else being
touched.
2026-08-29 20:32:18 +02:00
jschoubben ba31eef80f Point the example bundle at the current control plane 2026-08-29 19:58:02 +02:00
jschoubben 1bc97ed50d A service can be declared to reflect a file
Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.

That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.

Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.

Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.

Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
2026-08-29 18:04:16 +02:00
jschoubben 38d7b2d8af Point the example bundle at the current control plane image 2026-08-29 16:51:56 +02:00
jschoubben fa48b5825e The bundle and the mesh stop removing each other
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.

State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.

Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.

Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
2026-08-29 16:43:46 +02:00
jschoubben a488c76b5e A node holds its link open, and applies what the mesh signs
The loop the whole thing exists for: told, apply, report.

`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.

Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.

A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.

Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.

That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.
2026-08-29 16:23:27 +02:00
jschoubben a4445f5c0a A machine joins the mesh it raised
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.

Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.

The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.

An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.

Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.

The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.

Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.
2026-08-29 16:03:15 +02:00
jschoubben a740959cb0 The bootstrap reaches the broker, and survives a reboot
Five steps now instead of three. A sealed machine goes from bare to a container
runtime, a store, the inventory database, that database's schema applied by
mesh-control, and LavinMQ running and answering.

The database is called inventory rather than mesh. ADR 0008 grants a context
only what it exclusively owns and ADR 0006 says the mesh database names a thing
that will not exist -- so one database per context, and there is one context.

The broker is in the bundle because ADR 0006 now says it must be: the control
plane reaches a node only over the link, the link is the broker, so nothing can
provision the broker. Two images, which is the cost that record accepts.

Verified by reading the system rather than the report: inventory present and
mesh absent, the node table with its indexes, the migration row, lavinmqctl
answering, 5672 listening. Sealed confirmed both ways -- the internet times out,
the lab registry returns 200.

Then rebooted, which was the part worth doing rather than assuming. Everything
returned: docker from boot: enabled, both containers because this host creates
every container --restart unless-stopped, the schema intact in its volume. Three
reconciles before and one after all report no change.

One thing that reads as a success and was not: the first sealed check said the
machine could reach example.com. It was the test that was wrong -- a helper
script pasted arguments into a shell line, so a command with quotes was re-split
and ran on the workstation. The machine had been sealed the whole time. The
helper now requotes each argument.
2026-08-29 12:06:54 +02:00
jschoubben e09503acc7 A real bundle: a bare machine raises a store and a database
The first three steps of the substrate bootstrap, run on a lab machine
confirmed to have no route out. It went from bare to a container runtime
installed and enabled, PostgreSQL running from an image pinned by digest, and
the control plane's database created inside it -- from the file the host
carries, with nothing to ask.

Second run changed nothing. `owned` lists all five afterwards, and `mesh` is in
the store.

It stops before the last two steps because there is no control plane yet: its
schema cannot be loaded and its image does not exist. The bundle says so rather
than naming something that cannot be applied.

Two things fixed on the way.

`make host BUNDLE=...` still swapped a file called substrate.lock, which the
per-system split had renamed months of decisions ago -- it now takes SYSTEM and
replaces that system's bundle. And the .lock files still cited ADR 0060, since
the renumbering pass only covered .md, .go, .ts and .sh.

One thing learned by it failing first: a directory the host creates is owned by
root, and a database inside a container runs as somebody else, so it could not
write and the container crash-looped. The store's data is a named volume now,
which lets the image set up its own ownership and outlives the container --
which is what you want for the thing holding the mesh's state.

Worth noting the failure was caught by the action's verify rather than by the
container step. `docker inspect` reported the container running because it was,
briefly, between restarts. Running is not working, and the thing that knew the
difference was the step that asked the database whether it would answer.
2026-08-29 01:54:10 +02:00