Commit Graph
234 Commits
Author SHA1 Message Date
jochen bdd44154cc Judge how a module says it is ready, beside whether it stays up (hq ADR 0240, to-be 48 Phase B)
Liveness alone could not see a web application whose port was open and whose
program ran while every request hung for eleven hours (issue 145). A resource
now carries the `health` its module declared: the engine makes http and tcp
looks itself from the machine to the endpoint's published port, reads a unit's
readiness from the show it already makes, hands an exec command or the image's
own check to the runtime as the container's check with the declared timing and
reads its state from the inspect it already makes, and asks a module's tool on
its own node tools. Starting until the check passed, unhealthy once its looks
after the grace fail the declared number of times; never more looks than the
measured budget; nothing restarted. The statement says contract 2, which tells
the controller this engine may be sent the field.
2026-10-07 14:08:32 +02:00
jochen 3c1ac6aef2 Let a planned maintenance window fail no apply (hq issue 291)
mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
At 03:30 the store's collector held the registry still while the
node-engine's reconcile on that machine was fetching bundle blobs from
it: every archive failed 'connection refused' and the machine was held
until the next pass. Other machines can meet the same window.

A scheduled step now opens its window only once no apply is in flight
here (the apply lock is taken just to write the record, so a push still
never queues behind the window). An apply whose fetch the store does
not answer waits for a window open on its own machine to close, and
elsewhere retries with backoff within one bounded budget per apply,
well past the window's length; an answer such as 404 still fails at
once.
2026-10-07 13:11:03 +02:00
jochen 3a117c2d2b Keep a replaced build reachable, and delete it only once no process runs from it (hq issue 289)
mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
The witness moved a running controller's build into a 0700 directory and deleted it
on proof, while the old process could still be serving. Its directories are now
0711, and a build without a reader is retired and swept once /proc shows nothing
runs from it.
2026-10-07 02:45:08 +02:00
jochen 1fc1e74cc3 Judge whether what a module runs stays up, and say it (hq ADR 0240, to-be 48 Phase A)
mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
mesh/delivery-group group feat/a-module-says-how-it-is-healthy delivered: every member is delivered
A container that crash-looped after its compose applied passed every check the
gate had: nothing looked at what a module runs. The node-engine now judges every
long-running resource on every look — one read of the runtime, one per service
manager — keeps the restarts it counts across recreates and its own restarts,
and says the state in every report and as an event on change, again every minute
while not healthy. It reads only; nothing is restarted for being unhealthy.
2026-10-07 02:28:15 +02:00
jochen 2a4b521e1b Cite the hq issues by the numbers they were given: 285, 286, 287
mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (0 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery superseded: a newer head of the same pull request
2026-10-07 02:00:13 +02:00
jochen ed43d65bf3 Lay out the publishing test's return as every gofmt agrees, so the build seat's check passes
mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (0 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery superseded: a newer head of the same pull request
The toolchain the build seat runs merge-check.sh in (Go 1.26) and a newer local Go format a return of
several multi-line composite literals differently; the seat's is the one that judges, and it failed every
pull request on this file (novox/hq issue 283).
2026-10-07 01:29:06 +02:00
jochen 0c405b70cc Roll a core build back by a witness that is not the new build (hq to-be 45 Phase 4)
The launcher trusted a counter only a by-hand reconcile ever cleared and a
known-good nothing in the daemon wrote, so no machine could roll its host back;
the controller and the node tools were replaced in place with nothing kept.

- The launcher runs a delivered host that is not known-good on trial: one that
  crashes, stops for nothing, or does not report within ten minutes goes back
  to known-good, once per version, recorded in rolled-back. The host proves
  itself when the mesh takes a report under its own build, says every standing
  verdict on its reports, never stands aside for a rolled-back version, and
  restarts its service once when its launcher was replaced on disk.
- The engine keeps the controller's and the node tools' previous build beside
  the new one and judges the new one: the lease taken by the controller it
  started (read-only direct get of mesh-controller_lease/holder), or this
  machine's runtime answering $SRV.PING.node-tools.<node>, within sixty seconds
  of time it could ask. Not healthy: the previous restored, once, said. Proved:
  the previous deleted. A build declared not-reversible is never rolled back.
- Retire never removes a version newer than the running one.
2026-10-06 18:23:56 +02:00
jochen 7e9ae9f07e Carry the controller's new self-check grant; stop telling people to delete kept data (hq ADR 0233)
The installer's first user list must match the controller's composed grant, which now reads every
machine's backup holder; and a kept directory's report pointed at the one act that loses data.
2026-10-06 16:47:49 +02:00
jochen 98885e953e Genesis assigns mesh-wireguard, not the retired networking bundle (hq ADR 0226)
A controller that no longer ships networking refuses it, which would stop genesis where the hub
joins the private network.
2026-10-06 14:59:26 +02:00
jochen 6e1bcdbde4 Say again what was last applied when the mesh asks (hq to-be 45 Phase 3, the report verb)
The controller's healer H1 answers a send that went unreported by asking the
machine first: a report lost on its way (issue 264) needs no second send. The
node-engine now hears mesh.node.<self>.ask.report on core NATS and enqueues a
reconcile whose account is said whether or not it is news; a delivery waiting
meanwhile is applied and reported instead. The answer is an ordinary report on
its own subject, so the node publishes nothing new and answers nobody's inbox.

The genesis lock grants the controller the healer-acted seat event, which it
now composes; the genesis test in mesh-controller holds the two equal.
2026-10-06 14:05:02 +02:00
jochen 31804bd8c2 Apply through one queue, and order what is applied and reported (hq to-be 45 Phase 2)
A delivery and the five-minute reconcile were two paths that applied, ordered only by a lock, and
each order it allowed was met live (issues 257, 261, 267). Now both only enqueue: one worker takes
the newest declaration held when it starts, applies it once and makes one report, and reports leave
in the order they are made.

A declaration may carry the controller's lease epoch beside its sequence; one older than what this
node applied is refused before anything is touched, counted, logged and reported. A report carries
the declaration's epoch and sequence and the host's own report sequence, kept on disk so it goes on
increasing across restarts and self-updates. Without an epoch, today's behaviour stands.
2026-10-06 11:54:10 +02:00
jochen 6953b5bafd Say the heartbeat's interval, export the host's validator, grant the genesis controller Phase 1 (hq to-be 45)
The controller's watchdog of a machine's heartbeat (S1) is bound to three
of its intervals, and a bound the controller guessed would not move when
the interval does: the heartbeat now carries interval_seconds.

Its self-check (D1) must judge every composed declaration as the host
does, and a second validator written from the host's rules would drift
from them: the host's own parsing is exported, unchanged, as
github.com/novox/mesh-host/validate.

The installer's first user list grants what the controller now composes
for itself: its condition buckets, the condition and doctor-heartbeat
events, the bus's two consumer advisories and $SRV.INFO — or the first
controller would be refused them until the broker's machine is pushed.
2026-10-06 10:18:54 +02:00
jochen a79972577c Never say a reconcile's report after a newer apply's (hq issue 267)
A reconcile that held the machine to the kept declaration just before a
delivery arrived queued its report while the delivery's apply waited for
it; the link published the apply's report and then the reconcile's, so the
mesh's last word from the machine named the older declaration and the
release plan waited on a report it had already been given. The link now
sets aside an unasked report about a declaration other than the one it
has applied since.
2026-10-06 01:45:34 +02:00
jochen f62ee0bc75 Say an apply's report even when the apply ends the link
A host that delivered its own successor stood aside before the report of
that apply was published, so it failed with 'context canceled' and the
release plan waited for a report that never came (novox/hq issue 264).
Reports are now published on a context the stand-aside does not cancel,
and the last apply's report is kept until the broker takes it and said
again on the next link, so a crash between apply and report is covered too.
2026-10-06 00:30:15 +02:00
jochen 762da9e05e Hand a whole file to its new owner instead of removing it first (hq ADR 0223)
The resolver file moves from resolv-conf to the uplink's holder in one apply.
Orphans go first, so the file was deleted or given back its pre-mesh original
until the new owner's turn came. Now the old record is forgotten once the new
one is recorded at the same path, and the kept original goes with it.
2026-10-05 23:35:40 +02:00
jochen f0f5873202 Keep unitKey's comment on unitKey 2026-10-05 22:27:32 +02:00
jochen c9001abdb4 Leave a unit alone when another declared service still holds it (hq issue 190)
Removing one record of a unit gave back what that record found, even while another declared
service holds the unit: when the private network's docker.service record goes and the docker
module's stays, a record that found the runtime stopped would stop it, and every container with
it, only for the docker module to start it again in the same apply. The record is forgotten
instead, and the plan says so. A test pins the registry member moving from the network's record
to the docker module's in one apply without leaving the list.
2026-10-05 22:21:21 +02:00
jochen 101f0c0d61 Restart a service only after every file it names is written (hq issue 260)
A service was restarted where it was declared, so one declared ahead of a
file in its restart-on was restarted before that file existed (the
resolver's zones file, a fact appended after its module's resources), and
a later change to such a file was never acted on. Each service is now
applied right after the last resource it names under restart-on or
reload-on; nothing else moves.
2026-10-05 21:55:47 +02:00
jochen f08681f225 An apply leaves a container a maintenance window holds still (hq issue 224)
A while-stopped step stops its module's containers, and an apply arriving
mid-window read the stopped server as broken and recreated it running under
the step - for the store's collector, a registry taking an upload the sweep
then deletes. The scheduler now records each window under the node's state
directory before its first stop and erases it after its last start, so every
apply on the machine (daemon, reconcile by hand, installer) reports a held
container held-still and leaves it for the first apply after the window.
A window whose host died, or past six hours, holds nothing; the report says
which windows are open.
2026-10-05 18:27:04 +02:00
jschoubben b774a0b4f9 An undeclared file is deleted only when it holds what the host wrote; otherwise moved aside (hq issue 241) 2026-10-05 01:30:42 +02:00
jochen 3e11d720b4 A file written over is given its kept original back when undeclared (hq ADR 0118, 0102)
The host keeps the original of a file before writing over it (ADR 0102), but
removing the file's record deleted the file and never put the original back,
although ADR 0118 and the comment on meshMadeUnits say it does. A module writing
/etc/pacman.conf, logrotate.conf, locale.conf or vconsole.conf whole would, once
unassigned, leave the machine without the file.

removeWhole now decides, in order: no kept original (the mesh made it) is
removed as before; a file gone since is not brought back; a file changed since
the mesh last wrote it is left as it stands, as a block or JSON write-into stays
the machine's; an unreadable kept copy leaves the mesh's file in place. Otherwise
the original goes back atomically with the mode and owner it was found with,
now recorded beside Kept, and the outcome is "restored". None of it is fatal.
The plan says "restore" for such a file.

A kept original is carried only for the path it was kept from, and a file whose
path moved keeps the original at its new path first, so a moved file is never
given another path's original.
2026-10-04 12:54:16 +02:00
jochen d5c365cb2b A user unit waits for its account's manager, and lingering is the account's (hq ADR 0177)
An account's manager runs only while it is logged in or lingers. A user-scoped unit
whose manager is not running is now "waiting" rather than failed, its record kept as
it was; its removal is never fatal (kept recorded, retried) and an account that is
gone is forgotten. Whether the manager runs is asked of user@<uid>.service in the
machine's manager: asking the account's own, through --machine, logs it in.

The user shape gains `linger`, set with loginctl, read back from logind's record,
and given back on removal like the shell. Unit files the mesh writes under
~/.config/systemd/user or /etc/systemd/user make that unit the mesh's, and made,
holds and found units are keyed by manager and name, so an account's unit and the
machine's of one name are two units. A service moved between managers gives the old
one back through the manager it was in. OpenRC refuses both.
2026-10-04 12:41:32 +02:00
jochen 84540e709a A unit may be user-scoped: applied through the account's own manager (hq ADR 0177)
A workstation's per-user daemons — a window manager's reload watcher, an
audio mask, a memory guard — are units in the operator account's own service
manager, and until now had no form the mesh could send (to-be 29). The
`service` shape gains `scope` ("system", the default, or "user") and `user`
(the account, named ${machine:account} by a module); a user-scoped unit
without an account, or a system unit naming one, is refused at parse.

The host reaches the account's manager as `systemctl --user --machine=<account>@`
from its own process: no environment to forge, no user to switch to. Done on
the runner rather than per system, since every system's reading of a unit
already goes through systemctl. Apply, reflect-only and removal all go through
the same manager, and the applied record carries scope and user so removal
gives the unit back to the manager it came from. It answers only while that
manager runs — a login, or lingering enabled for the account; declaring
lingering is a follow-up.

Tests: a user-scoped unit is started and enabled in the account's manager and
recorded with its scope; a system unit never sees --user; the validation of
scope and user.
2026-10-04 12:31:24 +02:00
jochen 8470dbd8e8 An archive can be undeclared, and undeclaring one no longer stops the apply (hq issue 162)
An archive had no removal, so an unassigned one failed as an orphan and
aborted every apply after: a module with tools could not be unassigned,
and a race between two pushes froze a machine against every change.

The record now keeps what an archive unpacked: its files, the
directories the host made inside its path, whether the host made the
path itself, and the parents it made to reach it. Removal takes exactly
that away, directories only once empty, never one that was there
before; a directory that is the host's alone is renamed aside first so a
reader sees the whole bundle or none of it. Whatever cannot be removed
is said and forgotten, never fatal.

A directory found before the archive is no longer swapped away with
what was in it: the archive is moved in file by file, and one that would
write over a file the mesh did not put there is refused before anything
moves. A record from before this change learns its files from the
archive's bytes on the next apply; one already orphaned is left in
place, said and forgotten. A former target is still left in place: the
version before is what a rollback starts (ADR 0141).
2026-10-04 12:20:41 +02:00
jochen f2eda240ec Cite hq issue 228: 225 was taken on main while this branch was open 2026-10-04 10:30:57 +02:00
jochen 2a5f4c8270 Parents the host makes inside an owner's home are the owner's (hq ADR 0182, to-be 41)
A file or archive placed under a fresh account's home with an owner left
the parents it created, such as ~/.config or ~/.local/share, owned by
root, so the person's own programs could not write there. Parents that
already existed, and any outside the owner's home, are left as before.
2026-10-04 04:07:40 +02:00
jochen 5d5dccdd55 A login the mesh set is given back, and undeclaring one no longer stops the apply (hq issue 225)
A user had no removal, so an undeclared one failed as an orphan and
aborted every apply after. Removal now keeps the account, gives back
the shell recorded when the mesh first changed it if it is still the
mesh's and still usable, and says why otherwise (hq ADR 0176 §2).
A shell is refused before it is set unless it is executable and listed
in /etc/shells, since usermod succeeds on a missing one.
2026-10-04 03:57:34 +02:00
jschoubben b602b223e1 A changed maintenance window is a changed spec (hq ADR 0189)
containerSpec says a changed cadence moves the marker so the install is
reported updated and re-established. Which containers are held still is the
same kind of statement, and a declaration that changed it while the machine
reported no change would be a machine quietly holding yesterday's containers.
2026-10-04 03:27:32 +02:00
jschoubben e8c4824ae2 A scheduled step may hold its module's own containers still (hq ADR 0189)
while-stopped names resource ids of the same module's containers; the host
stops them before the run and starts them again after it, in reverse order,
whatever the step did. The restart is deferred before the first stop and runs
on its own context, because the one real risk of this field is a window that
never closes.

Scheduled steps only: at apply the declaration is applied in order and a
run-once step already gates what follows.
2026-10-04 02:34:33 +02:00
jochen 8280a82ef8 Read getent's exit code, not its wording (novox/hq issue 213)
A user that does not exist yet was matched as Go's 'exit status 2', while the host's runner says
'getent exited 2', so it read as a user database that did not answer: the controller's account was
never created and the handover to its process stopped there. The runner keeps the exit underneath
its words, and a caller asks the code.
2026-10-04 02:11:43 +02:00
jochen d9ea387680 Raise a process-form controller at genesis as the container it replaces (hq issue 223)
The controller's manifest now declares a Go bundle the host runs as a
process (novox/hq issue 213). Genesis cannot run that: the bundle is
fetched from the artifact store and compiled in a toolchain, and the mesh
makes both long after the controller. The builder, asked to build the
manifest at genesis, refuses for lack of the Go toolchain. So genesis
raises the controller as before, as a container, and the first push hands
it over to the process through `replaces` (issue 223, option b).

- Step 3 clones the controller at the commit with the carried builder's
  git and reads its manifest. In the image form (an older controller) it
  builds through the builder as before. In the process form it builds the
  repository's own Dockerfile and hands step 9 a manifest of its own
  shape: the process becomes a container with the id the process
  `replaces`, the image genesis built, host network, and every host path
  the process's env names mounted at that same path read-only. Secrets
  belong to the image's user (65534) until the process's account takes
  them over. `prepares` is dropped: the temporary controller from the
  same commit already migrated the stores, and a pinned image is
  nothing the controller can derive a step from.
- Steps 4 to 9 are unchanged: they take the manifest as they did.
- apply.ForTests lets the bootstrap's test apply a process.

The first composed declaration from the process manifest names
`mesh-controller.server`, which is what the host recorded for the genesis
container, so the first apply hands over and leaves one controller.
2026-10-04 01:47:33 +02:00
jochen 2ff3b50a84 Hand a replaced resource over to the process that replaces it (hq issue 213)
The controller moves from a container to a process on the one machine
that runs it (novox/hq issue 213). Every orphan is removed before anything
is applied, so the container would go first and nothing would answer the
mesh's verbs while the process was fetched, unpacked and started — and
never again, if it did not start.

- a process may say what it `replaces`: resources the declaration no
  longer declares. Such an orphan is kept through the up-front sweep and
  removed right after the process applied and is up: active and running
  at two looks ten seconds apart, the same main process, no restart in
  between (stricter than ADR 0184's second look, which reads a unit
  waiting to restart as running). If the process failed, was skipped
  behind its module's step, or is not running, the orphan stays running
  and recorded, reported kept, and the next apply hands it over.
  Refused: naming something still declared, itself, an empty id, one
  thing named by two processes, and `replaces` on a step or a schedule.
- beyond #85's oneshot unit for a step: a step written ./name runs its
  own bundle's binary (tested), and is started, never enabled.
- a run-once process that fails gates its module, as a run-once
  container already did, so a version whose preparation failed is not
  started.
- an unchanged run-once process is not run again, and an unchanged
  scheduled one is kept up by its timer: both were "a daemon that had
  stopped" and were started on every apply.
2026-10-04 01:01:52 +02:00
jochen 5fc4052a2b Run a run-once process as its oneshot unit (novox/hq design 38 WP4c)
A step was run directly: in the host's own working directory, without its env, env files or
user. A module step moved out of its container (node bootstrap/index.js) could find neither its
code nor its words. A oneshot unit carries all four as a daemon's does, and starting it waits.
2026-10-04 00:29:03 +02:00
jochen 1c29168309 Make an unpacked archive exactly the archive (novox/hq issue 220)
Unpacked over the previous tree, a file the new archive no longer has stayed: a bundle rebuilt as
one file per entrypoint kept the old package directory. Unpack into a fresh directory and swap it
in, so a refused archive also leaves the old tree whole.
2026-10-04 00:17:50 +02:00
jochen b7ede83c8c A process runs its own bundle's binary, written ./name (hq ADR 0193)
A bundle compiled to a binary runs itself, but only the host knows where it unpacked it, and the
service manager takes no path relative to the working directory. A command written ./name is that
file in the process's own unpacked bundle, made absolute in the unit.
2026-10-03 21:13:55 +02:00
jochen 482d20737f An unchanged process keeps its record, so the node's runtime is not re-created every other cycle (hq issue 210)
The process applier's unchanged path returned an outcome that said nothing about what was
written; the loop recorded it like any other, erasing the digest. The next cycle found no
record and re-created the daemon, the one after found a record again, and so on: the
node's runtime restarted every ten minutes on every machine since it arrived. The outcome
now carries the digest forward, as a file's does. The test applies one process three
times and asserts the record survives an unchanged apply and no restart is asked.
2026-10-03 13:41:14 +02:00
jochen c9b963f8b9 A failed install says whether the package database is stale, how old it is, and what fixes it (hq issue 205)
pacman writes its errors to stderr, which the runner folds into the error rather than the output;
the stale-index classifier read the output alone and never saw a single 'failed retrieving file',
so a ten-week-old database on the control node reported as a wall of 404s from the mirrors. The
classifier now reads everything pacman said, knows a failed signature as the same staleness, tells a
mirror outage with a fresh database apart from it, and names the database's date and age beside the
remedy: a full upgrade by the operator, never a one-package sync, which on this distribution is a
partial upgrade. Whose job keeping the database current is stays issue 205's question.
2026-10-03 03:54:17 +02:00
jschoubben b94e0f9a77 The mesh's own ban chain is a ban wherever it hangs; the front end's record is cited as 0180 (hq ADR 0186)
The legacy reader required every path into a chain of refusals to come from a built-in whose policy
accepts. The home server's ban chain hangs off the container runtime's user chain, whose forward
policy the runtime set to DROP, so the machine reported the mesh's own intrusion prevention as a
rule set the mesh did not write. A chain is a ban when every refusal names its sources and the
chain accepts nothing — the rule the nftables side already used. A chain that accepts anything is
still not a ban. Fixture captured from the machine. The citations for the uninstalled front end move
to ADR 0180, which another session's renumber had left pointing at an unrelated record.
2026-10-02 18:42:16 +02:00
jschoubben 286865dfa7 A service the mesh asked to run is still running a moment later (hq ADR 0184)
The read-back raced the failure: a service manager returns when it has started the process, and a
daemon that refuses its configuration exits a fraction of a second later, so one look saw it alive.
fail2ban took 221ms on the control node and the apply reported "restarted" onto a dead daemon while
both public machines kept no bans at all. The host looks again, after that moment. A unit still
starting is accepted at both looks; a service asked to stop is not waited on.
2026-10-02 18:16:24 +02:00
jschoubben b30d9c5b5a A container may log to the journal (hq ADR 0179)
A jail reads a log; a container's output went to a file of the runtime's own under a path that
changes on recreate, so no jail could read a container's service. logging: journald runs the
container with the journal as its driver, named in the spec so moving it recreates it; any other
place is refused.
2026-10-02 17:02:49 +02:00
jschoubben f57386cdea A package may be declared absent, and an uninstalled front end is retired for good (hq ADR 0175)
absent: true on a package has the host remove it through the machine's own
package manager when it is installed and leave alone a machine that never had
it; read back either way. Undeclaring a package still removes nothing. A
found firewall whose command is gone is recorded as removed, said once, and
asked nothing of.
2026-10-02 16:28:27 +02:00
jschoubben cdbe3ab0a4 Cite hq ADR 0170, not 0169: the firewall seat's record was renumbered after a collision on hq main 2026-10-02 14:52:20 +02:00
jschoubben ecb3003ba4 A machine reports whether its virtualisation daemon runs
The lab module needs the virtualisation daemon (novox/hq ADR 0172); the
capability is detected by asking the daemon about itself, not by finding
a client on disk.
2026-10-02 14:46:18 +02:00
jschoubben b6dbe0a7b9 A container may declare the capabilities it is granted (hq ADR 0169)
Exactly the names declared reach the runtime, named in the spec so a change
recreates the container; a name that is not a capability's is refused and a
privileged container stays undeclarable. For a seat holder whose runtime
changes the machine's packet filter.
2026-10-02 13:27:34 +02:00
jschoubben 627ac97d4f The host says what filters the machine, with owners, and keeps the found firewall retired on every converged apply (hq ADR 0168)
Every table and chain that refuses traffic is reported with whose it is:
the mesh's, the found firewall's, the container runtime's own, a ban, or
other — the runtime's user chain is other, which is where both predecessors
kept their rules, in the legacy filter on one machine and invisible to the
mesh. Adoption's threshold does not move; a converged machine's report
grows by its filters and its found firewall's state.

Convergence is a state the host keeps: a found firewall enabled again is
retired again and said; a reconcile that finds it inactive records that it
was found so, never that the mesh did it; a step skipped after a failed
apply is said. A retirement the mesh began and did not finish is finished.

Fixtures are rulesets captured from three machines of the first mesh.
2026-10-02 11:58:16 +02:00
jschoubben cbdbf6b7d3 Every physical link faces outside, up or down
The filter accepts what does not arrive on a link the machine names as
outward, and the host named only links carrying a default route. An
unplugged wired port was left unfiltered for whenever it was plugged in
(novox/hq issue 197). A link backed by a physical device is now named
whether or not it is up.
2026-10-02 11:53:53 +02:00
jschoubben c47aa5d9b3 A former target of a kind the host cannot remove is left in place and said, never fatal (hq issue 194)
The host delivers its own successor as an archive whose target is a new
directory each version, and since mesh-host 63 the record keeps a resource's
former target for the next apply to remove. An archive has no removal (issue
162), so the first host that replaced itself under that rule refused its own
former version at the first step of every apply, and all four machines applied
nothing from then on. A former target nobody dropped is forgotten and said;
an archive the declaration dropped still refuses.
2026-10-02 02:43:47 +02:00
jschoubben fb9c9c3ee8 A taken container keeps a found network, a left-out module is kept, and genesis raises the forge as its module declares (hq ADR 0163)
A container may name networks it also joins once created, for the per-machine
setting that keeps a found network while a neighbour still resolves it there:
joined after the run, part of the spec, refused when it cannot be joined.

A declaration may say which modules the mesh left out because a stored setting
cannot compose with its definition. Absence used to read as removal; a left-out
module's records are kept and said, and its holds are not released.

Genesis raises the bootstrap forge under the gitea module's container name, with
its image digest and its data directory mounted at /data, so the module holds it
by the found rule instead of raising a second forge beside it (issue 090). The
network is the one difference left for a take to say. Before this the forge had
no volume: its repositories were the container's, lost with it.
2026-10-01 23:45:05 +02:00
jschoubben 83b3d20e68 A take is a comparison: the host's facts, former targets, and strays (hq ADR 0163)
Every held thing carries what a take compares: for a found container its image and the image's date,
the networks it is on and the other containers on each, its mounts and published ports, beside the
declared image (and its date once pulled), ports and volumes, with the downgrade decided when both
dates are known; for a found file whether the declared content differs and how, as lines lost and
lines new. A resource whose target moved keeps the former target on record as an orphan, so the next
apply removes the container or file the host wrote under the old name (issue 097). Every apply reports
the strays: containers the mesh neither wrote nor holds.
2026-10-01 21:24:37 +02:00
jschoubben b1e9ccff6d The profile names the network manager that is running, and travels in every report (hq ADR 0161)
One capability per manager — uplink-networkmanager, uplink-systemd-networkd, uplink-dhcpcd — from
systemctl is-active, so the uplink seat's holder for a manager this machine does not run is refused
the way any missing capability is, naming it (issue 138). The apply that reports detects the profile
again and sends it, the same shape enrolment sends, so a machine that switched managers reaches the
mesh at its next push.
2026-10-01 15:58:31 +02:00