The host keeps the original of a file before writing over it (ADR 0102), but
removing the file's record deleted the file and never put the original back,
although ADR 0118 and the comment on meshMadeUnits say it does. A module writing
/etc/pacman.conf, logrotate.conf, locale.conf or vconsole.conf whole would, once
unassigned, leave the machine without the file.
removeWhole now decides, in order: no kept original (the mesh made it) is
removed as before; a file gone since is not brought back; a file changed since
the mesh last wrote it is left as it stands, as a block or JSON write-into stays
the machine's; an unreadable kept copy leaves the mesh's file in place. Otherwise
the original goes back atomically with the mode and owner it was found with,
now recorded beside Kept, and the outcome is "restored". None of it is fatal.
The plan says "restore" for such a file.
A kept original is carried only for the path it was kept from, and a file whose
path moved keeps the original at its new path first, so a moved file is never
given another path's original.
An account's manager runs only while it is logged in or lingers. A user-scoped unit
whose manager is not running is now "waiting" rather than failed, its record kept as
it was; its removal is never fatal (kept recorded, retried) and an account that is
gone is forgotten. Whether the manager runs is asked of user@<uid>.service in the
machine's manager: asking the account's own, through --machine, logs it in.
The user shape gains `linger`, set with loginctl, read back from logind's record,
and given back on removal like the shell. Unit files the mesh writes under
~/.config/systemd/user or /etc/systemd/user make that unit the mesh's, and made,
holds and found units are keyed by manager and name, so an account's unit and the
machine's of one name are two units. A service moved between managers gives the old
one back through the manager it was in. OpenRC refuses both.
A workstation's per-user daemons — a window manager's reload watcher, an
audio mask, a memory guard — are units in the operator account's own service
manager, and until now had no form the mesh could send (to-be 29). The
`service` shape gains `scope` ("system", the default, or "user") and `user`
(the account, named ${machine:account} by a module); a user-scoped unit
without an account, or a system unit naming one, is refused at parse.
The host reaches the account's manager as `systemctl --user --machine=<account>@`
from its own process: no environment to forge, no user to switch to. Done on
the runner rather than per system, since every system's reading of a unit
already goes through systemctl. Apply, reflect-only and removal all go through
the same manager, and the applied record carries scope and user so removal
gives the unit back to the manager it came from. It answers only while that
manager runs — a login, or lingering enabled for the account; declaring
lingering is a follow-up.
Tests: a user-scoped unit is started and enabled in the account's manager and
recorded with its scope; a system unit never sees --user; the validation of
scope and user.
An archive had no removal, so an unassigned one failed as an orphan and
aborted every apply after: a module with tools could not be unassigned,
and a race between two pushes froze a machine against every change.
The record now keeps what an archive unpacked: its files, the
directories the host made inside its path, whether the host made the
path itself, and the parents it made to reach it. Removal takes exactly
that away, directories only once empty, never one that was there
before; a directory that is the host's alone is renamed aside first so a
reader sees the whole bundle or none of it. Whatever cannot be removed
is said and forgotten, never fatal.
A directory found before the archive is no longer swapped away with
what was in it: the archive is moved in file by file, and one that would
write over a file the mesh did not put there is refused before anything
moves. A record from before this change learns its files from the
archive's bytes on the next apply; one already orphaned is left in
place, said and forgotten. A former target is still left in place: the
version before is what a rollback starts (ADR 0141).
A file or archive placed under a fresh account's home with an owner left
the parents it created, such as ~/.config or ~/.local/share, owned by
root, so the person's own programs could not write there. Parents that
already existed, and any outside the owner's home, are left as before.
A user had no removal, so an undeclared one failed as an orphan and
aborted every apply after. Removal now keeps the account, gives back
the shell recorded when the mesh first changed it if it is still the
mesh's and still usable, and says why otherwise (hq ADR 0176 §2).
A shell is refused before it is set unless it is executable and listed
in /etc/shells, since usermod succeeds on a missing one.
A user that does not exist yet was matched as Go's 'exit status 2', while the host's runner says
'getent exited 2', so it read as a user database that did not answer: the controller's account was
never created and the handover to its process stopped there. The runner keeps the exit underneath
its words, and a caller asks the code.
The controller's manifest now declares a Go bundle the host runs as a
process (novox/hq issue 213). Genesis cannot run that: the bundle is
fetched from the artifact store and compiled in a toolchain, and the mesh
makes both long after the controller. The builder, asked to build the
manifest at genesis, refuses for lack of the Go toolchain. So genesis
raises the controller as before, as a container, and the first push hands
it over to the process through `replaces` (issue 223, option b).
- Step 3 clones the controller at the commit with the carried builder's
git and reads its manifest. In the image form (an older controller) it
builds through the builder as before. In the process form it builds the
repository's own Dockerfile and hands step 9 a manifest of its own
shape: the process becomes a container with the id the process
`replaces`, the image genesis built, host network, and every host path
the process's env names mounted at that same path read-only. Secrets
belong to the image's user (65534) until the process's account takes
them over. `prepares` is dropped: the temporary controller from the
same commit already migrated the stores, and a pinned image is
nothing the controller can derive a step from.
- Steps 4 to 9 are unchanged: they take the manifest as they did.
- apply.ForTests lets the bootstrap's test apply a process.
The first composed declaration from the process manifest names
`mesh-controller.server`, which is what the host recorded for the genesis
container, so the first apply hands over and leaves one controller.
The controller moves from a container to a process on the one machine
that runs it (novox/hq issue 213). Every orphan is removed before anything
is applied, so the container would go first and nothing would answer the
mesh's verbs while the process was fetched, unpacked and started — and
never again, if it did not start.
- a process may say what it `replaces`: resources the declaration no
longer declares. Such an orphan is kept through the up-front sweep and
removed right after the process applied and is up: active and running
at two looks ten seconds apart, the same main process, no restart in
between (stricter than ADR 0184's second look, which reads a unit
waiting to restart as running). If the process failed, was skipped
behind its module's step, or is not running, the orphan stays running
and recorded, reported kept, and the next apply hands it over.
Refused: naming something still declared, itself, an empty id, one
thing named by two processes, and `replaces` on a step or a schedule.
- beyond #85's oneshot unit for a step: a step written ./name runs its
own bundle's binary (tested), and is started, never enabled.
- a run-once process that fails gates its module, as a run-once
container already did, so a version whose preparation failed is not
started.
- an unchanged run-once process is not run again, and an unchanged
scheduled one is kept up by its timer: both were "a daemon that had
stopped" and were started on every apply.
A step was run directly: in the host's own working directory, without its env, env files or
user. A module step moved out of its container (node bootstrap/index.js) could find neither its
code nor its words. A oneshot unit carries all four as a daemon's does, and starting it waits.
Unpacked over the previous tree, a file the new archive no longer has stayed: a bundle rebuilt as
one file per entrypoint kept the old package directory. Unpack into a fresh directory and swap it
in, so a refused archive also leaves the old tree whole.
The controller announces the mesh-controller seat on $SRV.PING/$SRV.INFO from genesis; the
carried account is the composed one, which the controller's test compares.
A bundle compiled to a binary runs itself, but only the host knows where it unpacked it, and the
service manager takes no path relative to the working directory. A command written ./name is that
file in the process's own unpacked bundle, made absolute in the unit.
The process applier's unchanged path returned an outcome that said nothing about what was
written; the loop recorded it like any other, erasing the digest. The next cycle found no
record and re-created the daemon, the one after found a record again, and so on: the
node's runtime restarted every ten minutes on every machine since it arrived. The outcome
now carries the digest forward, as a file's does. The test applies one process three
times and asserts the record survives an unchanged apply and no restart is asked.
pacman writes its errors to stderr, which the runner folds into the error rather than the output;
the stale-index classifier read the output alone and never saw a single 'failed retrieving file',
so a ten-week-old database on the control node reported as a wall of 404s from the mirrors. The
classifier now reads everything pacman said, knows a failed signature as the same staleness, tells a
mirror outage with a fresh database apart from it, and names the database's date and age beside the
remedy: a full upgrade by the operator, never a one-package sync, which on this distribution is a
partial upgrade. Whose job keeping the database current is stays issue 205's question.
Until the retired mesh-build-machine row is dropped the controller asks and hears both roles, so
the genesis list grants both; the controller's own test of this list says so.
The controller composes the build seat's subjects as node-build-agent's now; the installer's
genesis list must say the same, or the controller's own test of that list disagrees with it.
Put back by hand, it was found afresh and its copy became the hold's original, so a machine that
kept restoring it kept growing copies and lost which one was first. The first original now stays
the record's, content that differs is kept once beside it, and the note says a rollback means
unassigning the private network. Nothing is removed unless it is <wireguard dir>/<iface>.conf,
not a path the mesh writes, and not a link, which would leave the key-bearing target behind.
Kept on disk it was the take's fallback; once the mesh's interface is up in its place and a peer
has handshaken with it, it is an unmaintained way back onto the network, held for ever. It is now
removed from where its unit reads it, its kept original verified first and left as it is, and the
hold ends. Until proven — no handshake, or wg not answering — it is kept and the report says why.
The retirement is recorded apart from holds, so later applies, an undeclare, and a reassignment
find it retired rather than missing, and nothing writes it back.
Records written before Found existed left the adoption guard and the converge filter loaded on
undeclare, then deleted their unit files from under them; a unit whose own file the mesh created
is now the mesh's, whatever its record says. Found is kept apart the moment it is read, so a
first apply that enabled and then failed is not read back as the machine's; boot is found the
first time the mesh sets it; a service once stateless, or moved to another unit, is found afresh
(the old unit given back). The unit is read after the reload that loads a file written in the
same apply, and removal reports what it actually did.
removeProcess deletes filepath.Join(daemonRoot, name) whole; a process named ".." made that
/var/lib/mesh. The declaration and the removal now hold the name to one rule.
A service undeclared used to be stopped: unassigning the private network stopped the container
runtime, unassigning sshd would stop ssh, an uplink module would take the machine offline. The
host now records the unit's state when it first applies it and restores that on undeclare —
found running stays running; started by the mesh (the converge filter) is stopped again; nothing
is started on the way out; a pre-existing record leaves the unit alone.
An undeclared process had no removal at all and failed every apply on its node; its unit, timer
and bundle are now removed.