Compare commits

...
55 Commits
Author SHA1 Message Date
mesh-admin 89fce1dae3 Merge pull request 'A file the host wrote over is given back when undeclared (hq ADR 0102, 0118)' (#92) from fix/a-file-written-over-is-given-back into main 2026-10-04 10:54:50 +00:00
jochen 3e11d720b4 A file written over is given its kept original back when undeclared (hq ADR 0118, 0102)
The host keeps the original of a file before writing over it (ADR 0102), but
removing the file's record deleted the file and never put the original back,
although ADR 0118 and the comment on meshMadeUnits say it does. A module writing
/etc/pacman.conf, logrotate.conf, locale.conf or vconsole.conf whole would, once
unassigned, leave the machine without the file.

removeWhole now decides, in order: no kept original (the mesh made it) is
removed as before; a file gone since is not brought back; a file changed since
the mesh last wrote it is left as it stands, as a block or JSON write-into stays
the machine's; an unreadable kept copy leaves the mesh's file in place. Otherwise
the original goes back atomically with the mode and owner it was found with,
now recorded beside Kept, and the outcome is "restored". None of it is fatal.
The plan says "restore" for such a file.

A kept original is carried only for the path it was kept from, and a file whose
path moved keeps the original at its new path first, so a moved file is never
given another path's original.
2026-10-04 12:54:16 +02:00
mesh-admin 429672b7ff Merge pull request 'User-scoped units (hq ADR 0177, to-be 38 WP6), with lingering and a manager that is not running' (#91) from feat/user-scoped-units into main 2026-10-04 10:43:17 +00:00
jochen d5c365cb2b A user unit waits for its account's manager, and lingering is the account's (hq ADR 0177)
An account's manager runs only while it is logged in or lingers. A user-scoped unit
whose manager is not running is now "waiting" rather than failed, its record kept as
it was; its removal is never fatal (kept recorded, retried) and an account that is
gone is forgotten. Whether the manager runs is asked of user@<uid>.service in the
machine's manager: asking the account's own, through --machine, logs it in.

The user shape gains `linger`, set with loginctl, read back from logind's record,
and given back on removal like the shell. Unit files the mesh writes under
~/.config/systemd/user or /etc/systemd/user make that unit the mesh's, and made,
holds and found units are keyed by manager and name, so an account's unit and the
machine's of one name are two units. A service moved between managers gives the old
one back through the manager it was in. OpenRC refuses both.
2026-10-04 12:41:32 +02:00
jochen 84540e709a A unit may be user-scoped: applied through the account's own manager (hq ADR 0177)
A workstation's per-user daemons — a window manager's reload watcher, an
audio mask, a memory guard — are units in the operator account's own service
manager, and until now had no form the mesh could send (to-be 29). The
`service` shape gains `scope` ("system", the default, or "user") and `user`
(the account, named ${machine:account} by a module); a user-scoped unit
without an account, or a system unit naming one, is refused at parse.

The host reaches the account's manager as `systemctl --user --machine=<account>@`
from its own process: no environment to forge, no user to switch to. Done on
the runner rather than per system, since every system's reading of a unit
already goes through systemctl. Apply, reflect-only and removal all go through
the same manager, and the applied record carries scope and user so removal
gives the unit back to the manager it came from. It answers only while that
manager runs — a login, or lingering enabled for the account; declaring
lingering is a follow-up.

Tests: a user-scoped unit is started and enabled in the account's manager and
recorded with its scope; a system unit never sees --user; the validation of
scope and user.
2026-10-04 12:31:24 +02:00
mesh-admin 6431848342 Merge pull request 'An archive can be undeclared, and undeclaring one no longer stops the apply (hq issue 162)' (#90) from fix/162-an-archive-can-be-undeclared into main 2026-10-04 10:22:39 +00:00
jochen 8470dbd8e8 An archive can be undeclared, and undeclaring one no longer stops the apply (hq issue 162)
An archive had no removal, so an unassigned one failed as an orphan and
aborted every apply after: a module with tools could not be unassigned,
and a race between two pushes froze a machine against every change.

The record now keeps what an archive unpacked: its files, the
directories the host made inside its path, whether the host made the
path itself, and the parents it made to reach it. Removal takes exactly
that away, directories only once empty, never one that was there
before; a directory that is the host's alone is renamed aside first so a
reader sees the whole bundle or none of it. Whatever cannot be removed
is said and forgotten, never fatal.

A directory found before the archive is no longer swapped away with
what was in it: the archive is moved in file by file, and one that would
write over a file the mesh did not put there is refused before anything
moves. A record from before this change learns its files from the
archive's bytes on the next apply; one already orphaned is left in
place, said and forgotten. A former target is still left in place: the
version before is what a rollback starts (ADR 0141).
2026-10-04 12:20:41 +02:00
mesh-admin 64b421b6e1 Merge pull request 'A login the mesh set is given back, and directories made inside a home are its account's (hq issue 228, to-be 41 WP1)' (#89) from feat/the-shell-and-its-environment into main 2026-10-04 08:49:30 +00:00
jochen f2eda240ec Cite hq issue 228: 225 was taken on main while this branch was open 2026-10-04 10:30:57 +02:00
jochen 2a5f4c8270 Parents the host makes inside an owner's home are the owner's (hq ADR 0182, to-be 41)
A file or archive placed under a fresh account's home with an owner left
the parents it created, such as ~/.config or ~/.local/share, owned by
root, so the person's own programs could not write there. Parents that
already existed, and any outside the owner's home, are left as before.
2026-10-04 04:07:40 +02:00
jochen 5d5dccdd55 A login the mesh set is given back, and undeclaring one no longer stops the apply (hq issue 225)
A user had no removal, so an undeclared one failed as an orphan and
aborted every apply after. Removal now keeps the account, gives back
the shell recorded when the mesh first changed it if it is still the
mesh's and still usable, and says why otherwise (hq ADR 0176 §2).
A shell is refused before it is set unless it is executable and listed
in /etc/shells, since usermod succeeds on a missing one.
2026-10-04 03:57:34 +02:00
mesh-admin 27fb22fddb Merge pull request 'A scheduled step may hold its module's own containers still (hq ADR 0189, issue 108)' (#77) from feat/the-store-keeps-what-the-records-name into main 2026-10-04 01:39:07 +00:00
jschoubben b602b223e1 A changed maintenance window is a changed spec (hq ADR 0189)
containerSpec says a changed cadence moves the marker so the install is
reported updated and re-established. Which containers are held still is the
same kind of statement, and a declaration that changed it while the machine
reported no change would be a machine quietly holding yesterday's containers.
2026-10-04 03:27:32 +02:00
jschoubben e8c4824ae2 A scheduled step may hold its module's own containers still (hq ADR 0189)
while-stopped names resource ids of the same module's containers; the host
stops them before the run and starts them again after it, in reverse order,
whatever the step did. The restart is deferred before the first stop and runs
on its own context, because the one real risk of this field is a window that
never closes.

Scheduled steps only: at apply the declaration is applied in order and a
run-once step already gates what follows.
2026-10-04 02:34:33 +02:00
mesh-admin 8c2e76f4c1 Merge pull request 'Read getent's exit code, not its wording (hq issue 213)' (#88) from fix/getent-not-found-is-an-answer into main 2026-10-04 00:11:48 +00:00
jochen 8280a82ef8 Read getent's exit code, not its wording (novox/hq issue 213)
A user that does not exist yet was matched as Go's 'exit status 2', while the host's runner says
'getent exited 2', so it read as a user database that did not answer: the controller's account was
never created and the handover to its process stopped there. The runner keeps the exit underneath
its words, and a caller asks the code.
2026-10-04 02:11:43 +02:00
mesh-admin 2d5e76434b Merge pull request 'Raise a process-form controller at genesis as the container it replaces (hq issue 223)' (#87) from fix/issue-223-genesis-pivots-to-the-controllers-container into main 2026-10-03 23:50:48 +00:00
jochen d9ea387680 Raise a process-form controller at genesis as the container it replaces (hq issue 223)
The controller's manifest now declares a Go bundle the host runs as a
process (novox/hq issue 213). Genesis cannot run that: the bundle is
fetched from the artifact store and compiled in a toolchain, and the mesh
makes both long after the controller. The builder, asked to build the
manifest at genesis, refuses for lack of the Go toolchain. So genesis
raises the controller as before, as a container, and the first push hands
it over to the process through `replaces` (issue 223, option b).

- Step 3 clones the controller at the commit with the carried builder's
  git and reads its manifest. In the image form (an older controller) it
  builds through the builder as before. In the process form it builds the
  repository's own Dockerfile and hands step 9 a manifest of its own
  shape: the process becomes a container with the id the process
  `replaces`, the image genesis built, host network, and every host path
  the process's env names mounted at that same path read-only. Secrets
  belong to the image's user (65534) until the process's account takes
  them over. `prepares` is dropped: the temporary controller from the
  same commit already migrated the stores, and a pinned image is
  nothing the controller can derive a step from.
- Steps 4 to 9 are unchanged: they take the manifest as they did.
- apply.ForTests lets the bootstrap's test apply a process.

The first composed declaration from the process manifest names
`mesh-controller.server`, which is what the host recorded for the genesis
container, so the first apply hands over and leaves one controller.
2026-10-04 01:47:33 +02:00
mesh-admin 99ad145bec Merge pull request 'Hand a replaced resource over to the process that replaces it (hq issue 213)' (#86) from fix/issue-213-the-controller-is-a-process into main 2026-10-03 23:13:52 +00:00
jochen 2ff3b50a84 Hand a replaced resource over to the process that replaces it (hq issue 213)
The controller moves from a container to a process on the one machine
that runs it (novox/hq issue 213). Every orphan is removed before anything
is applied, so the container would go first and nothing would answer the
mesh's verbs while the process was fetched, unpacked and started — and
never again, if it did not start.

- a process may say what it `replaces`: resources the declaration no
  longer declares. Such an orphan is kept through the up-front sweep and
  removed right after the process applied and is up: active and running
  at two looks ten seconds apart, the same main process, no restart in
  between (stricter than ADR 0184's second look, which reads a unit
  waiting to restart as running). If the process failed, was skipped
  behind its module's step, or is not running, the orphan stays running
  and recorded, reported kept, and the next apply hands it over.
  Refused: naming something still declared, itself, an empty id, one
  thing named by two processes, and `replaces` on a step or a schedule.
- beyond #85's oneshot unit for a step: a step written ./name runs its
  own bundle's binary (tested), and is started, never enabled.
- a run-once process that fails gates its module, as a run-once
  container already did, so a version whose preparation failed is not
  started.
- an unchanged run-once process is not run again, and an unchanged
  scheduled one is kept up by its timer: both were "a daemon that had
  stopped" and were started on every apply.
2026-10-04 01:01:52 +02:00
mesh-admin a24670d77c Merge pull request 'Run a run-once process as its oneshot unit (design 38 WP4c)' (#85) from fix/a-run-once-step-runs-where-and-as-declared into main 2026-10-03 22:29:10 +00:00
jochen 5fc4052a2b Run a run-once process as its oneshot unit (novox/hq design 38 WP4c)
A step was run directly: in the host's own working directory, without its env, env files or
user. A module step moved out of its container (node bootstrap/index.js) could find neither its
code nor its words. A oneshot unit carries all four as a daemon's does, and starting it waits.
2026-10-04 00:29:03 +02:00
mesh-admin a7bf0f6e39 Merge pull request 'Make an unpacked archive exactly the archive (hq issue 220)' (#84) from fix/issue-220-a-bundle-on-disk-is-exactly-the-artifact into main 2026-10-03 22:17:57 +00:00
jochen 1c29168309 Make an unpacked archive exactly the archive (novox/hq issue 220)
Unpacked over the previous tree, a file the new archive no longer has stayed: a bundle rebuilt as
one file per entrypoint kept the old package directory. Unpack into a fresh directory and swap it
in, so a refused archive also leaves the old tree whole.
2026-10-04 00:17:50 +02:00
mesh-admin 483e02e7ef Merge pull request 'The installer's first user list lets the controller answer $SRV.STATS (hq ADR 0197)' (#83) from fix/0197-the-controller-answers-stats-too into main 2026-10-03 21:11:25 +00:00
jochen 0650df9414 The installer's first user list lets the controller answer $SRV.STATS (hq ADR 0197) 2026-10-03 22:47:41 +02:00
mesh-admin 47fb924947 Merge pull request 'The installer's first user list lets the controller answer discovery for its seat (hq ADR 0197)' (#82) from feat/0197-the-controller-announces-itself into main 2026-10-03 20:11:12 +00:00
jochen 09f84b451d The installer's first user list lets the controller answer discovery for its seat (hq ADR 0197)
The controller announces the mesh-controller seat on $SRV.PING/$SRV.INFO from genesis; the
carried account is the composed one, which the controller's test compares.
2026-10-03 22:11:01 +02:00
mesh-admin 77ae5320ad Merge pull request 'A process runs its own bundle's binary, written ./name (hq ADR 0193)' (#81) from feat/a-process-runs-its-own-binary into main 2026-10-03 19:14:04 +00:00
jochen b7ede83c8c A process runs its own bundle's binary, written ./name (hq ADR 0193)
A bundle compiled to a binary runs itself, but only the host knows where it unpacked it, and the
service manager takes no path relative to the working directory. A command written ./name is that
file in the process's own unpacked bundle, made absolute in the unit.
2026-10-03 21:13:55 +02:00
mesh-admin 195fc63f6a Merge pull request 'An unchanged process keeps its record, so the node's runtime is not re-created every other cycle (hq issue 210)' (#80) from fix/issue-210-a-process-is-unchanged-the-second-time into main 2026-10-03 11:41:34 +00:00
jochen 482d20737f An unchanged process keeps its record, so the node's runtime is not re-created every other cycle (hq issue 210)
The process applier's unchanged path returned an outcome that said nothing about what was
written; the loop recorded it like any other, erasing the digest. The next cycle found no
record and re-created the daemon, the one after found a record again, and so on: the
node's runtime restarted every ten minutes on every machine since it arrived. The outcome
now carries the digest forward, as a file's does. The test applies one process three
times and asserts the record survives an unchanged apply and no restart is asked.
2026-10-03 13:41:14 +02:00
mesh-admin 09614d94c8 Merge pull request 'A failed install says whether the package database is stale, how old it is, and what fixes it (hq issue 205)' (#79) from fix/issue-205 into main 2026-10-03 09:44:43 +00:00
jochen c9b963f8b9 A failed install says whether the package database is stale, how old it is, and what fixes it (hq issue 205)
pacman writes its errors to stderr, which the runner folds into the error rather than the output;
the stale-index classifier read the output alone and never saw a single 'failed retrieving file',
so a ten-week-old database on the control node reported as a wall of 404s from the mirrors. The
classifier now reads everything pacman said, knows a failed signature as the same staleness, tells a
mirror outage with a fresh database apart from it, and names the database's date and age beside the
remedy: a full upgrade by the operator, never a one-package sync, which on this distribution is a
partial upgrade. Whose job keeping the database current is stays issue 205's question.
2026-10-03 03:54:17 +02:00
mesh-admin 5a6e51f975 Merge pull request 'The first user list names the build role as node-build-agent (hq ADR 0190)' (#78) from fix/the-build-role-is-node-build-agent into main 2026-10-03 00:53:44 +00:00
jochen 3c96683d98 The first user list carries both build seats during the handover (hq ADR 0190)
Until the retired mesh-build-machine row is dropped the controller asks and hears both roles, so
the genesis list grants both; the controller's own test of this list says so.
2026-10-03 02:53:12 +02:00
jochen 8c4e33c165 The first user list names the build role as node-build-agent (hq ADR 0190)
The controller composes the build seat's subjects as node-build-agent's now; the installer's
genesis list must say the same, or the controller's own test of that list disagrees with it.
2026-10-02 22:37:08 +02:00
mesh-admin 2a332b0e6e Merge pull request 'The mesh's own ban chain is a ban wherever it hangs (hq ADR 0186)' (#76) from fix/a-ban-list-never-holds-a-neighbour into main 2026-10-02 16:43:09 +00:00
jschoubben b94e0f9a77 The mesh's own ban chain is a ban wherever it hangs; the front end's record is cited as 0180 (hq ADR 0186)
The legacy reader required every path into a chain of refusals to come from a built-in whose policy
accepts. The home server's ban chain hangs off the container runtime's user chain, whose forward
policy the runtime set to DROP, so the machine reported the mesh's own intrusion prevention as a
rule set the mesh did not write. A chain is a ban when every refusal names its sources and the
chain accepts nothing — the rule the nftables side already used. A chain that accepts anything is
still not a ban. Fixture captured from the machine. The citations for the uninstalled front end move
to ADR 0180, which another session's renumber had left pointing at an unrelated record.
2026-10-02 18:42:16 +02:00
mesh-admin b840ce0a77 Merge pull request 'A service the mesh asked to run is still running a moment later (hq ADR 0184)' (#75) from fix/a-service-asked-to-run-is-still-running into main 2026-10-02 16:25:25 +00:00
jschoubben 286865dfa7 A service the mesh asked to run is still running a moment later (hq ADR 0184)
The read-back raced the failure: a service manager returns when it has started the process, and a
daemon that refuses its configuration exits a fraction of a second later, so one look saw it alive.
fail2ban took 221ms on the control node and the apply reported "restarted" onto a dead daemon while
both public machines kept no bans at all. The host looks again, after that moment. A unit still
starting is accepted at both looks; a service asked to stop is not waited on.
2026-10-02 18:16:24 +02:00
mesh-admin ca7c4a5915 Merge pull request 'A container may log to the journal (hq ADR 0179)' (#73) from feat/the-intrusion-seat-serves-its-verbs into main 2026-10-02 15:04:02 +00:00
jschoubben b30d9c5b5a A container may log to the journal (hq ADR 0179)
A jail reads a log; a container's output went to a file of the runtime's own under a path that
changes on recreate, so no jail could read a container's service. logging: journald runs the
container with the journal as its driver, named in the spec so moving it recreates it; any other
place is refused.
2026-10-02 17:02:49 +02:00
mesh-admin 5b73e04192 Merge pull request 'A package may be declared absent, and an uninstalled front end is retired for good (hq ADR 0175)' (#71) from feat/the-found-front-end-is-uninstalled into main 2026-10-02 14:29:17 +00:00
jschoubben f57386cdea A package may be declared absent, and an uninstalled front end is retired for good (hq ADR 0175)
absent: true on a package has the host remove it through the machine's own
package manager when it is installed and leave alone a machine that never had
it; read back either way. Undeclaring a package still removes nothing. A
found firewall whose command is gone is recorded as removed, said once, and
asked nothing of.
2026-10-02 16:28:27 +02:00
mesh-admin f92dd3e286 Merge pull request 'Cite hq ADR 0170, not 0169: the firewall seat's record was renumbered' (#70) from fix/adr-0170-cited into main 2026-10-02 12:53:09 +00:00
jschoubben cdbe3ab0a4 Cite hq ADR 0170, not 0169: the firewall seat's record was renumbered after a collision on hq main 2026-10-02 14:52:20 +02:00
mesh-admin a8f1cdb445 Merge pull request 'A machine reports whether its virtualisation daemon runs (hq ADR 0172)' (#69) from jschoubben/the-lab-is-a-module into main 2026-10-02 12:47:55 +00:00
jschoubben ecb3003ba4 A machine reports whether its virtualisation daemon runs
The lab module needs the virtualisation daemon (novox/hq ADR 0172); the
capability is detected by asking the daemon about itself, not by finding
a client on disk.
2026-10-02 14:46:18 +02:00
mesh-admin d3861f82d4 Merge pull request 'A container may declare the capabilities it is granted (hq ADR 0169)' (#68) from feat/the-firewall-seat-serves-its-verbs into main 2026-10-02 11:28:34 +00:00
jschoubben b6dbe0a7b9 A container may declare the capabilities it is granted (hq ADR 0169)
Exactly the names declared reach the runtime, named in the spec so a change
recreates the container; a name that is not a capability's is refused and a
privileged container stays undeclarable. For a seat holder whose runtime
changes the machine's packet filter.
2026-10-02 13:27:34 +02:00
mesh-admin 07bdad9e94 Merge pull request 'The host says what filters the machine, with owners, and keeps the found firewall retired on every converged apply (hq ADR 0168)' (#67) from feat/one-thing-filters-a-converged-machine into main 2026-10-02 09:58:58 +00:00
jschoubben 627ac97d4f The host says what filters the machine, with owners, and keeps the found firewall retired on every converged apply (hq ADR 0168)
Every table and chain that refuses traffic is reported with whose it is:
the mesh's, the found firewall's, the container runtime's own, a ban, or
other — the runtime's user chain is other, which is where both predecessors
kept their rules, in the legacy filter on one machine and invisible to the
mesh. Adoption's threshold does not move; a converged machine's report
grows by its filters and its found firewall's state.

Convergence is a state the host keeps: a found firewall enabled again is
retired again and said; a reconcile that finds it inactive records that it
was found so, never that the mesh did it; a step skipped after a failed
apply is said. A retirement the mesh began and did not finish is finished.

Fixtures are rulesets captured from three machines of the first mesh.
2026-10-02 11:58:16 +02:00
mesh-admin ee2359f29f Merge pull request 'Every physical link faces outside, up or down (hq issue 197)' (#66) from jschoubben/every-physical-link-faces-outside into main 2026-10-02 09:54:06 +00:00
jschoubben cbdbf6b7d3 Every physical link faces outside, up or down
The filter accepts what does not arrive on a link the machine names as
outward, and the host named only links carrying a default route. An
unplugged wired port was left unfiltered for whenever it was plugged in
(novox/hq issue 197). A link backed by a physical device is now named
whether or not it is up.
2026-10-02 11:53:53 +02:00
61 changed files with 7857 additions and 253 deletions
+33 -2
View File
@@ -1141,6 +1141,16 @@ func adoptionFingerprint(r link.Report) string {
for _, h := range r.Held { for _, h := range r.Held {
parts = append(parts, "held "+h.ID+"="+h.Changed) parts = append(parts, "held "+h.ID+"="+h.Changed)
} }
// And what filters the machine, with the found firewall's state (novox/hq ADR 0168): a rule the
// operator removes between declarations, or a front end enabled again, is said at the next
// reconcile rather than at the next push.
for _, f := range r.Filters {
parts = append(parts, "filter "+f.Owner+" "+f.Where+" "+f.Refuses)
}
if r.FoundFirewall != nil {
parts = append(parts, fmt.Sprintf("found-firewall %s active=%v retired-by=%s", r.FoundFirewall.Kind,
r.FoundFirewall.Active, r.FoundFirewall.RetiredBy))
}
for _, reach := range r.Reachable { for _, reach := range r.Reachable {
parts = append(parts, fmt.Sprintf("reach %s %s:%d %s %v %d", reach.Protocol, reach.Address, parts = append(parts, fmt.Sprintf("reach %s %s:%d %s %v %d", reach.Protocol, reach.Address,
reach.Port, reach.By, reach.Published, reach.ContainerPort)) reach.Port, reach.By, reach.Published, reach.ContainerPort))
@@ -1289,7 +1299,8 @@ func worthSaying(report link.Report) bool {
if report.Refused != "" { if report.Refused != "" {
return false return false
} }
return len(report.Held) > 0 || report.Firewall != "" || len(report.Outward) > 0 return len(report.Held) > 0 || report.Firewall != "" || len(report.Outward) > 0 ||
len(report.Filters) > 0 || report.FoundFirewall != nil
} }
// applyDeclared applies a declaration that has already been proved to come from the mesh. // applyDeclared applies a declaration that has already been proved to come from the mesh.
@@ -1421,7 +1432,7 @@ func applyAndKeep(ctx context.Context, opts options, raw []byte, signed *store.D
// (novox/hq ADR 0140). Reported whatever the node's mode: a converged node's filter needs it, // (novox/hq ADR 0140). Reported whatever the node's mode: a converged node's filter needs it,
// and an adopted one becomes converged without a further round trip. A machine that cannot read // and an adopted one becomes converged without a further round trip. A machine that cannot read
// its own routing table says nothing rather than guessing, and is sent no filter. // its own routing table says nothing rather than guessing, and is sent no filter.
if links, err := outward.Links(""); err != nil { if links, err := outward.Links("", ""); err != nil {
fmt.Fprintf(os.Stderr, "mesh-host: applied, and could not read which links face outside: %v\n", err) fmt.Fprintf(os.Stderr, "mesh-host: applied, and could not read which links face outside: %v\n", err)
} else { } else {
report.Outward = links report.Outward = links
@@ -1440,6 +1451,21 @@ func applyAndKeep(ctx context.Context, opts options, raw []byte, signed *store.D
report.Strays = append(report.Strays, link.Stray{Kind: s.Kind, Name: s.Name, Detail: s.Detail}) report.Strays = append(report.Strays, link.Stray{Kind: s.Kind, Name: s.Name, Detail: s.Detail})
} }
} }
// What filters this machine, with owners, whatever its mode (novox/hq ADR 0168): the mesh says
// truthfully what filters a converged machine, and names what it did not write.
ufwActive := firewall.Active(ctx, apply.ExecRunner)
if filters, err := firewall.Collect(ctx, apply.ExecRunner, ufwActive); err != nil {
fmt.Fprintf(os.Stderr, "mesh-host: applied, and could not read what filters this machine: %v\n", err)
} else {
for _, f := range filters {
report.Filters = append(report.Filters, link.Filter{Where: f.Where, Owner: f.Owner, Refuses: f.Refuses})
}
}
if declared.Adoption == nil && updated.Firewall != nil && updated.Firewall.Kind == string(firewall.UFW) && updated.Firewall.WasActive {
// And, converged, the state of the firewall it was found with and who retired it.
report.FoundFirewall = &link.FoundFirewall{Kind: updated.Firewall.Kind, Active: ufwActive,
RetiredBy: updated.Firewall.RetiredBy}
}
if declared.Adoption != nil { if declared.Adoption != nil {
if updated.Firewall != nil { if updated.Firewall != nil {
report.Firewall = updated.Firewall.Kind report.Firewall = updated.Firewall.Kind
@@ -1461,6 +1487,11 @@ func applyAndKeep(ctx context.Context, opts options, raw []byte, signed *store.D
if change.Action == "held" { if change.Action == "held" {
continue continue
} }
// Nor is a unit waiting for its account's manager that this machine never applied (novox/hq
// ADR 0177); one applied before and waiting now is still the machine's, recorded as it was.
if _, recorded := updated.Find(change.ID); change.Action == "waiting" && !recorded {
continue
}
report.Applied = append(report.Applied, change.ID) report.Applied = append(report.Applied, change.ID)
} }
// Kept whichever way it went, so a node that is disconnected next minute still knows what it // Kept whichever way it went, so a node that is disconnected next minute still knows what it
+16
View File
@@ -171,6 +171,22 @@ func TestAReconcileSpeaksOnlyWhenWhatIsHeldChanged(t *testing.T) {
if !w.changed(link.Report{Firewall: "none", Held: rewritten.Held}) { if !w.changed(link.Report{Firewall: "none", Held: rewritten.Held}) {
t.Error("a changed firewall was not said") t.Error("a changed firewall was not said")
} }
// What filters the machine is part of it (novox/hq ADR 0168): a predecessor's chain removed by
// hand, or the found firewall enabled again, is said without being asked.
filtered := link.Report{Firewall: "none", Held: rewritten.Held,
Filters: []link.Filter{{Where: "chain HAL-MESH-ONLY (iptables-legacy)", Owner: "other", Refuses: "-j DROP"}}}
if !w.changed(filtered) {
t.Error("a filter appearing was not said")
}
if !w.changed(link.Report{Firewall: "none", Held: rewritten.Held}) {
t.Error("a filter removed was not said")
}
if !w.changed(link.Report{Firewall: "none", Held: rewritten.Held, FoundFirewall: &link.FoundFirewall{Kind: "ufw", Active: true}}) {
t.Error("the found firewall coming back was not said")
}
if !worthSaying(link.Report{Filters: filtered.Filters}) {
t.Error("a report carrying only what filters the machine is not worth saying")
}
} }
func TestWhatTheLinkPublishedCountsAsSaid(t *testing.T) { func TestWhatTheLinkPublishedCountsAsSaid(t *testing.T) {
+1 -1
View File
@@ -161,7 +161,7 @@
"type": "file", "type": "file",
"path": "/var/lib/mesh-bus-conf/accounts.conf", "path": "/var/lib/mesh-bus-conf/accounts.conf",
"mode": "0600", "mode": "0600",
"content": "// The first user list, carried by the installer because at genesis there is no mesh to\n// compose one. A bootstrap credential, rotated with the store's and replaced by the\n// controller's own composition from its first start onward.\naccounts {\n MESH {\n jetstream: enabled\n users = [\n { user: \"controller\", password: \"$2a$10$AHqJgOifIVbU41KmATiMhuXFs8xa7Wl2HuN4UVBCXdN2jIQzjqApy\", permissions: {\n publish: { allow: [\"$JS.ACK.CONTROL.controller.>\", \"$JS.ACK.EVENTS.controller.>\", \"$JS.API.>\", \"_INBOX.enrol.>\", \"mesh.assignment.>\", \"mesh.control.>\", \"mesh.mod.*.tool.>\", \"mesh.node.>\", \"mesh.seat.mesh-build-machine.accept.>\", \"mesh.seat.mesh-controller.event.applied\", \"mesh.seat.mesh-controller.event.built-before\", \"mesh.seat.mesh-controller.event.refused\"] }\n subscribe: { allow: [\"$JS.API.>\", \"_DELIVER.controller\", \"_DELIVER.controller.>\", \"_INBOX.controller.>\", \"mesh.control.>\", \"mesh.mod.gitea.event.pull.merged\", \"mesh.mod.mesh-catalog.event.catching-up\", \"mesh.mod.mesh-catalog.event.upgraded\", \"mesh.seat.mesh-build-machine.event.built\", \"mesh.seat.mesh-controller.tool.>\"] }\n allow_responses: { max: 1, ttl: \"1m\" }\n } }\n ]\n }\n}\n" "content": "// The first user list, carried by the installer because at genesis there is no mesh to\n// compose one. A bootstrap credential, rotated with the store's and replaced by the\n// controller's own composition from its first start onward.\naccounts {\n MESH {\n jetstream: enabled\n users = [\n { user: \"controller\", password: \"$2a$10$AHqJgOifIVbU41KmATiMhuXFs8xa7Wl2HuN4UVBCXdN2jIQzjqApy\", permissions: {\n publish: { allow: [\"$JS.ACK.CONTROL.controller.>\", \"$JS.ACK.EVENTS.controller.>\", \"$JS.API.>\", \"_INBOX.enrol.>\", \"mesh.assignment.>\", \"mesh.control.>\", \"mesh.mod.*.tool.>\", \"mesh.node.>\", \"mesh.seat.mesh-build-machine.accept.>\", \"mesh.seat.node-build-agent.accept.>\", \"mesh.seat.mesh-controller.event.applied\", \"mesh.seat.mesh-controller.event.built-before\", \"mesh.seat.mesh-controller.event.refused\"] }\n subscribe: { allow: [\"$JS.API.>\", \"_DELIVER.controller\", \"_DELIVER.controller.>\", \"_INBOX.controller.>\", \"mesh.control.>\", \"mesh.mod.gitea.event.pull.merged\", \"mesh.mod.mesh-catalog.event.catching-up\", \"mesh.mod.mesh-catalog.event.upgraded\", \"mesh.seat.mesh-build-machine.event.built\", \"mesh.seat.node-build-agent.event.built\", \"mesh.seat.mesh-controller.tool.>\", \"$SRV.PING\", \"$SRV.INFO\", \"$SRV.PING.mesh-controller\", \"$SRV.PING.mesh-controller.>\", \"$SRV.INFO.mesh-controller\", \"$SRV.INFO.mesh-controller.>\", \"$SRV.STATS\", \"$SRV.STATS.mesh-controller\", \"$SRV.STATS.mesh-controller.>\"] }\n allow_responses: { max: 1, ttl: \"1m\" }\n } }\n ]\n }\n}\n"
}, },
{ {
"id": "broker", "id": "broker",
+573 -45
View File
@@ -27,6 +27,7 @@ import (
"time" "time"
"github.com/novox/mesh-host/internal/declaration" "github.com/novox/mesh-host/internal/declaration"
"github.com/novox/mesh-host/internal/firewall"
"github.com/novox/mesh-host/internal/store" "github.com/novox/mesh-host/internal/store"
"github.com/novox/mesh-host/internal/system" "github.com/novox/mesh-host/internal/system"
) )
@@ -53,12 +54,23 @@ type Outcome struct {
wrote string wrote string
// into is what a file written into held before the mesh's keys (novox/hq ADR 0102). // into is what a file written into held before the mesh's keys (novox/hq ADR 0102).
into *store.Into into *store.Into
// kept is where this apply kept the original of a file it wrote over (novox/hq ADR 0100). // kept is where this apply kept the original of a file it wrote over (novox/hq ADR 0100), and
kept string // keptMode and keptOwner how the original was found, in the form a hold records them.
kept, keptMode, keptOwner string
// stateless is a service whose unit's lifecycle is the machine's (novox/hq ADR 0117). // stateless is a service whose unit's lifecycle is the machine's (novox/hq ADR 0117).
stateless bool stateless bool
// scope and user are, for a service, whose manager it was applied through (novox/hq ADR 0177).
scope, user string
// found is, for a service, its unit as the host first found it (novox/hq ADR 0118). // found is, for a service, its unit as the host first found it (novox/hq ADR 0118).
found *store.FoundUnit found *store.FoundUnit
// shell is, for a user, the login shell it was found with and the one the mesh set (novox/hq
// ADR 0176 §2, issue 228).
shell *store.LoginShell
// linger is, for a user, whether it lingered before the mesh changed that, and what the mesh
// set (novox/hq ADR 0177).
linger *store.Lingering
// unpacked is, for an archive, what it put on the machine (novox/hq issue 162).
unpacked *store.Unpacked
// reads is, for a container, the digest of each file it was created reading, by path — so // reads is, for a container, the digest of each file it was created reading, by path — so
// the next apply can say which one changed (novox/hq 04-ISSUES/103). // the next apply can say which one changed (novox/hq 04-ISSUES/103).
reads map[string]string reads map[string]string
@@ -67,6 +79,10 @@ type Outcome struct {
// Report is what an apply did, in the order it did it. // Report is what an apply did, in the order it did it.
type Report struct { type Report struct {
Outcomes []Outcome `json:"outcomes"` Outcomes []Outcome `json:"outcomes"`
// Firewall is what this apply did about the firewall a converged machine was found with, when
// it did or declined anything: retired, retired again, or left in force and why (novox/hq ADR
// 0168). Said rather than an outcome: the plan says the same step the same way.
Firewall string `json:"firewall,omitempty"`
// Tunnel is what this apply says about the tunnel the private network took over, when the // Tunnel is what this apply says about the tunnel the private network took over, when the
// declaration names one (novox/hq ADR 0105). // declaration names one (novox/hq ADR 0105).
Tunnel *TakenTunnel `json:"tunnel,omitempty"` Tunnel *TakenTunnel `json:"tunnel,omitempty"`
@@ -76,8 +92,9 @@ type Report struct {
// nothing is the ordinary steady state, and saying so is not the same as saying it failed. // nothing is the ordinary steady state, and saying so is not the same as saying it failed.
func (r Report) Changed() bool { func (r Report) Changed() bool {
for _, o := range r.Outcomes { for _, o := range r.Outcomes {
// Holding is keeping the machine as it was found, which is not moving it. // Holding is keeping the machine as it was found, which is not moving it; waiting is a unit
if o.Action != "unchanged" && o.Action != "held" { // whose account's manager is not running, which nothing moved either (novox/hq ADR 0177).
if o.Action != "unchanged" && o.Action != "held" && o.Action != "waiting" {
return true return true
} }
} }
@@ -201,12 +218,13 @@ func ApplyKeeping(
} }
if errors.Is(err, errNoRemoval) && store.IsFormer(orphan.ID) { if errors.Is(err, errNoRemoval) && store.IsFormer(orphan.ID) {
// **A former target of a kind the host cannot remove is left in place and forgotten, // **A former target of a kind the host cannot remove is left in place and forgotten,
// never fatal.** The host's own archive is the case: every version it delivers itself // never fatal.** The host's own archive was the case (novox/hq issue 194): every version
// has a new target, so the one before is a former target on the first apply of the new // it delivers itself has a new target, so the one before is a former target on the first
// host — and a removal that refused there stopped every machine applying anything, the // apply of the new host — and a removal that refused there stopped every machine applying
// moment the host that carried former targets (novox/hq ADR 0163, rule 5) first // anything, the moment the host that carried former targets (novox/hq ADR 0163, rule 5)
// replaced itself. What was written stays where it is, said, and the record no longer // first replaced itself. What was written stays where it is, said, and the record no
// names it; whether an archive gets a removal is issue 162's question, not this apply's. // longer names it. An archive answers this itself since issue 162 (removeArchive); this
// stays for any kind that still has no removal.
known.Forget(orphan.ID) known.Forget(orphan.ID)
report.Outcomes = append(report.Outcomes, Outcome{ report.Outcomes = append(report.Outcomes, Outcome{
ID: orphan.ID, Type: orphan.Type, Target: orphan.Target, ID: orphan.ID, Type: orphan.Type, Target: orphan.Target,
@@ -215,6 +233,15 @@ func ApplyKeeping(
log(fmt.Sprintf(" forgotten %s (%s): a former target left in place: %v", orphan.ID, orphan.Target, err)) log(fmt.Sprintf(" forgotten %s (%s): a former target left in place: %v", orphan.ID, orphan.Target, err))
return nil return nil
} }
if errors.Is(err, errRemovalWaits) {
// Kept recorded and said, never fatal: tried again by the next apply (novox/hq ADR 0177).
report.Outcomes = append(report.Outcomes, Outcome{
ID: orphan.ID, Type: orphan.Type, Target: orphan.Target,
Action: "waiting", Detail: err.Error(),
})
log(fmt.Sprintf(" waiting %s (%s): %v", orphan.ID, orphan.Target, err))
return nil
}
if err != nil { if err != nil {
return &Error{Resource: orphan.ID, Err: err, Done: report} return &Error{Resource: orphan.ID, Err: err, Done: report}
} }
@@ -242,6 +269,20 @@ func ApplyKeeping(
// orphan is removed, and if a removal then fails the guard is already up. A stale opening on an // orphan is removed, and if a removal then fails the guard is already up. A stale opening on an
// adopted node is removed as any orphan is. // adopted node is removed as any orphan is.
var protecting, orphans []store.Applied var protecting, orphans []store.Applied
// **What a declared process replaces is handed over, not removed first** (novox/hq issue 213).
// Every other orphan goes before anything is applied; one a process names under `replaces` is
// kept until that process is applied and still running a moment later, so whatever it was —
// the controller's container — answers until its replacement does, and goes on answering if
// the replacement never comes up.
replacedBy := map[string]string{}
for _, r := range d.Resources {
if p, ok := r.(*declaration.Process); ok {
for _, id := range p.Replaces {
replacedBy[id] = p.ID
}
}
}
handover := map[string][]store.Applied{}
for _, orphan := range known.Orphans(declared, origin) { for _, orphan := range known.Orphans(declared, origin) {
if d.Adoption == nil && strings.HasPrefix(orphan.ID, declaration.AdoptionPrefix) { if d.Adoption == nil && strings.HasPrefix(orphan.ID, declaration.AdoptionPrefix) {
protecting = append(protecting, orphan) protecting = append(protecting, orphan)
@@ -259,6 +300,10 @@ func ApplyKeeping(
log(fmt.Sprintf(" kept %s (%s): %s was left out of this declaration by the mesh, not removed", orphan.ID, orphan.Target, module)) log(fmt.Sprintf(" kept %s (%s): %s was left out of this declaration by the mesh, not removed", orphan.ID, orphan.Target, module))
continue continue
} }
if by, replaced := replacedBy[orphan.ID]; replaced {
handover[by] = append(handover[by], orphan)
continue
}
orphans = append(orphans, orphan) orphans = append(orphans, orphan)
} }
ordered := d.Resources ordered := d.Resources
@@ -461,7 +506,8 @@ func ApplyKeeping(
// A file this host has no record of, under any id, is the machine's until the mesh // A file this host has no record of, under any id, is the machine's until the mesh
// writes over it — on any node, adopted or not: its original is kept first. // writes over it — on any node, adopted or not: its original is kept first.
var keepFound Keep var keepFound Keep
if f, isFile := resource.(*declaration.File); isFile && was.ID == "" && // A file whose path moved is a file this host has no record of at its new path.
if f, isFile := resource.(*declaration.File); isFile && (was.ID == "" || was.Target != f.Path) &&
!known.Recorded(string(declaration.TypeFile), f.Path) { !known.Recorded(string(declaration.TypeFile), f.Path) {
keepFound = keep keepFound = keep
} }
@@ -512,6 +558,11 @@ func ApplyKeeping(
if c, ok := resource.(*declaration.Container); ok && c.RunOnce { if c, ok := resource.(*declaration.Container); ok && c.RunOnce {
gates = true gates = true
} }
// And a run-once process, which is the same step hosted as a unit: a version whose
// preparation did not complete must not be started (novox/hq ADR 0135, issue 213).
if p, ok := resource.(*declaration.Process); ok && p.RunOnce {
gates = true
}
if gates { if gates {
failed.Gated = true failed.Gated = true
// **A module's step gates that module, not the machine** (novox/hq ADR 0136). // **A module's step gates that module, not the machine** (novox/hq ADR 0136).
@@ -538,15 +589,31 @@ func ApplyKeeping(
continue continue
} }
// **Waiting is not applied** (novox/hq ADR 0177): a user-scoped unit whose account's manager
// is not running was not touched, so its record — if it has one — stays exactly as it was,
// what was found included, and none is written for one never applied. What one whose
// manager stopped part way found is kept as a failure's is, for the apply that finishes it.
if outcome.Action == "waiting" {
if outcome.found != nil {
known.KeepFound(resource.Identity(), origin, *outcome.found)
}
report.Outcomes = append(report.Outcomes, outcome)
log(fmt.Sprintf(" waiting %s (%s): %s", outcome.ID, outcome.Target, outcome.Detail))
continue
}
// Where the original of what this file replaced was kept, carried for as long as the // Where the original of what this file replaced was kept, carried for as long as the
// resource is recorded: kept by this apply, by a hold its module's cutover ends, or before. // resource is recorded: kept by this apply, by a hold its module's cutover ends, or before.
// Carried with how the original was found, and only for the path it was kept from: a file
// whose path moved leaves its original with the record of the old path (a former target),
// and must never be given another path's original when it goes (novox/hq ADR 0118).
held, wasHeld := known.HeldAt(resource.Identity()) held, wasHeld := known.HeldAt(resource.Identity())
kept := outcome.kept kept, keptMode, keptOwner := outcome.kept, outcome.keptMode, outcome.keptOwner
if kept == "" && wasHeld { if kept == "" && wasHeld && held.Target == outcome.Target {
kept = held.Kept kept, keptMode, keptOwner = held.Kept, held.Mode, held.Owner
} }
if kept == "" { if kept == "" && was.Target == outcome.Target {
kept = was.Kept kept, keptMode, keptOwner = was.Kept, was.KeptMode, was.KeptOwner
} }
// Only now. The record follows the fact, never leads it. // Only now. The record follows the fact, never leads it.
known.Record(store.Applied{ known.Record(store.Applied{
@@ -556,9 +623,16 @@ func ApplyKeeping(
Wrote: outcome.wrote, Wrote: outcome.wrote,
Into: outcome.into, Into: outcome.into,
Kept: kept, Kept: kept,
KeptMode: keptMode,
KeptOwner: keptOwner,
Reads: outcome.reads, Reads: outcome.reads,
Stateless: outcome.stateless, Stateless: outcome.stateless,
Scope: outcome.scope,
User: outcome.user,
Found: outcome.found, Found: outcome.found,
Shell: outcome.shell,
Linger: outcome.linger,
Unpacked: outcome.unpacked,
Holds: holds(resource), Holds: holds(resource),
}) })
if outcome.found != nil { if outcome.found != nil {
@@ -602,6 +676,28 @@ func ApplyKeeping(
} }
log(line) log(line)
} }
// Its replacement applied: what it replaces goes now, once it is seen running (issue 213).
if waiting, has := handover[resource.Identity()]; has {
delete(handover, resource.Identity())
if err := handOver(ctx, run, resource, waiting, removeOrphan, &report, log); err != nil {
failures = append(failures, err)
}
}
}
// A replacement that did not apply — failed, skipped behind its module's step, held — leaves
// what it replaces running and recorded, said, for the next apply to hand over.
unhanded := make([]string, 0, len(handover))
for by := range handover {
unhanded = append(unhanded, by)
}
sort.Strings(unhanded)
for _, by := range unhanded {
for _, orphan := range handover[by] {
keptForReplacement(&report, log, orphan,
fmt.Sprintf("kept: %s, which replaces it, did not apply", by))
}
} }
if !orphansRemoved { if !orphansRemoved {
@@ -611,16 +707,24 @@ func ApplyKeeping(
} }
// A converged node whose found firewall was in force retires it only now, once everything — // A converged node whose found firewall was in force retires it only now, once everything —
// the mesh's derived filter among it — applied cleanly (novox/hq ADR 0100). // the mesh's derived filter among it — applied cleanly (novox/hq ADR 0100), and on every
// converged apply, not once (ADR 0168). Skipped, it is said: a step that does nothing is never
// silent (issue 143).
if len(failures) == 0 { if len(failures) == 0 {
if err := retireFirewall(ctx, d, origin, &known, run, log); err != nil { did, err := retireFirewall(ctx, d, origin, &known, run, log)
if err != nil {
return report, known, &Error{Resource: "the firewall found on this machine", Err: err, Done: report} return report, known, &Error{Resource: "the firewall found on this machine", Err: err, Done: report}
} }
report.Firewall = did
for _, orphan := range protecting { for _, orphan := range protecting {
if err := removeOrphan(orphan); err != nil { if err := removeOrphan(orphan); err != nil {
return report, known, err return report, known, err
} }
} }
} else if rec := known.Firewall; origin == store.OriginDeclared && d.Adoption == nil && rec != nil &&
rec.Kind == string(firewall.UFW) && rec.WasActive && firewall.Active(ctx, run) {
report.Firewall = fmt.Sprintf("left in force: %d resource(s) failed, and the found firewall is retired only after a clean apply", len(failures))
log(" kept ufw in force: " + report.Firewall)
} }
if len(failures) > 0 { if len(failures) > 0 {
@@ -658,7 +762,7 @@ func applyOne(ctx context.Context, sys system.System, r declaration.Resource, ru
case *declaration.Container: case *declaration.Container:
return applyContainer(ctx, res, run, changed, in, previous) return applyContainer(ctx, res, run, changed, in, previous)
case *declaration.User: case *declaration.User:
return applyUser(ctx, sys, res, run) return applyUser(ctx, sys, res, run, previous)
case *declaration.Archive: case *declaration.Archive:
return applyArchive(ctx, res, previous) return applyArchive(ctx, res, previous)
case *declaration.Process: case *declaration.Process:
@@ -710,7 +814,7 @@ func applyDirectory(r *declaration.Directory) (Outcome, error) {
} }
if !existed { if !existed {
if err := os.MkdirAll(r.Path, mode); err != nil { if err := makeDirs(r.Path, mode, r.Owner); err != nil {
return out, err return out, err
} }
} }
@@ -883,9 +987,13 @@ func applyFile(r *declaration.File, previous store.Applied, unseal Unseal, keepF
} }
var beforeMode os.FileMode var beforeMode os.FileMode
beforeOwner := ""
if existed { if existed {
if info, err := os.Stat(r.Path); err == nil { if info, err := os.Stat(r.Path); err == nil {
beforeMode = info.Mode().Perm() beforeMode = info.Mode().Perm()
if uid, gid, ok := ownerOf(info); ok {
beforeOwner = fmt.Sprintf("%d:%d", uid, gid)
}
} }
} }
@@ -906,7 +1014,7 @@ func applyFile(r *declaration.File, previous store.Applied, unseal Unseal, keepF
return out, fmt.Errorf("keeping the original of %s before writing over it: %w", r.Path, err) return out, fmt.Errorf("keeping the original of %s before writing over it: %w", r.Path, err)
} }
} }
if err := os.MkdirAll(filepath.Dir(r.Path), 0o755); err != nil { if err := makeDirs(filepath.Dir(r.Path), 0o755, r.Owner); err != nil {
return out, err return out, err
} }
if err := writeAtomically(r.Path, []byte(content), mode); err != nil { if err := writeAtomically(r.Path, []byte(content), mode); err != nil {
@@ -971,6 +1079,7 @@ func applyFile(r *declaration.File, previous store.Applied, unseal Unseal, keepF
} }
out.kept = kept out.kept = kept
if kept != "" { if kept != "" {
out.keptMode, out.keptOwner = fmt.Sprintf("%04o", beforeMode), beforeOwner
if out.Detail != "" { if out.Detail != "" {
out.Detail += "; " out.Detail += "; "
} }
@@ -1028,12 +1137,93 @@ type unitReloader interface {
ReloadUnits(ctx context.Context, run system.Runner) error ReloadUnits(ctx context.Context, run system.Runner) error
} }
// serviceSettle is how long the host waits before looking at a unit a second time. A test sets it
// to nothing; on a machine it is the window in which a daemon that refuses its configuration dies.
var serviceSettle = 2 * time.Second
// stayedRunning is the state of a unit the host has just asked to run, read twice.
//
// **Because the first read races the failure.** A service manager returns when it has started the
// process, and the unit is "activating" or "active" at that instant whatever the process is about
// to do. A daemon that reads its configuration, refuses it and exits does so a fraction of a second
// later — fail2ban took 221 milliseconds the day this was written — so a single read back says
// running about a machine whose daemon is already gone, and the apply reports "restarted" for a
// service that is dead. Every ban on both public machines was lost that way while every check
// passed (novox/hq ADR 0184), which is the one shape of failure this host exists to refuse.
//
// So it looks again, after the moment in which that happens. It does not wait for a slow unit to
// finish starting: a unit still coming up reads as running both times and is accepted, as before.
// What this catches is a unit that was running and is not any more.
func stayedRunning(ctx context.Context, sys system.System, run Runner, unit string) (string, error) {
state, err := sys.ServiceState(ctx, run, unit)
if err != nil || state != "running" {
return state, err
}
timer := time.NewTimer(serviceSettle)
defer timer.Stop()
select {
case <-ctx.Done():
return state, ctx.Err()
case <-timer.C:
}
return sys.ServiceState(ctx, run, unit)
}
// applyService puts a unit into its declared state, in the manager it belongs to.
//
// **A user-scoped unit waits for its account's manager** (novox/hq ADR 0177). That manager runs
// from the account's first login to its last logout, or always while the account lingers; with
// neither, there is nothing to start the unit in, and asking would log the account in for the
// length of the question (see UserManagerRunning). A unit that fails every apply until somebody logs
// in is a node reported broken for being switched on, so it is said — "waiting", recorded as it
// was, nothing claimed — and the first apply after the manager starts applies it. One the manager
// itself starts at login needs nothing from the mesh: enabled is enabled in the account's own
// manager, and that starts what is enabled when it starts.
func applyService(ctx context.Context, sys system.System, r *declaration.Service, run Runner, func applyService(ctx context.Context, sys system.System, r *declaration.Service, run Runner,
changed map[string]bool, previous store.Applied) (Outcome, error) { changed map[string]bool, previous store.Applied) (Outcome, error) {
if !r.UserScoped() {
return applyServiceIn(ctx, sys, r, run, changed, previous)
}
away, err := managerAway(ctx, sys, r.Scope, r.User, run)
if err != nil {
return begin(r), err
}
if away != "" {
return waitingFor(r, away), nil
}
out, err := applyServiceIn(ctx, sys, r, run, changed, previous)
if err != nil {
// A manager that stopped while the apply was at it — the person logged out — is the same
// absence found a moment later, and is said the same way rather than failed.
// What was found before it went is carried, as a failure's is (novox/hq ADR 0118).
if away, _ := managerAway(ctx, sys, r.Scope, r.User, run); away != "" {
waiting := waitingFor(r, away+", having stopped during this apply ("+err.Error()+")")
waiting.found = out.found
return waiting, nil
}
}
return out, err
}
// waitingFor is a user-scoped unit whose account's manager is not running (novox/hq ADR 0177).
func waitingFor(r *declaration.Service, away string) Outcome {
out := begin(r)
out.scope, out.user = r.Scope, r.User
out.Action = "waiting"
out.Detail = away + "; applied by the first apply after it starts — a login, or the account " +
"declared to linger"
return out
}
// applyServiceIn is applyService, in the manager the unit belongs to; raw reaches the machine's.
func applyServiceIn(ctx context.Context, sys system.System, r *declaration.Service, raw Runner,
changed map[string]bool, previous store.Applied) (Outcome, error) {
run := managerFor(r.Scope, r.User, raw)
if r.Stateless() { if r.Stateless() {
return reflectOnly(ctx, sys, r, run, changed) return reflectOnly(ctx, sys, r, run, changed)
} }
out := begin(r) out := begin(r)
out.scope, out.user = r.Scope, r.User
var changes []string var changes []string
// A file the service reflects changed, and it may be the unit's own file or a drop-in: the // A file the service reflects changed, and it may be the unit's own file or a drop-in: the
@@ -1053,7 +1243,7 @@ func applyService(ctx context.Context, sys system.System, r *declaration.Service
// running, from the mesh's packet filter, stopped until the mesh started it and to be stopped // running, from the mesh's packet filter, stopped until the mesh started it and to be stopped
// again. Read after the reload above, never before it: a unit whose file this same apply wrote // again. Read after the reload above, never before it: a unit whose file this same apply wrote
// is not a unit the service manager knows until then, and reading it first would find nothing. // is not a unit the service manager knows until then, and reading it first would find nothing.
found, gaveBack, err := foundAs(ctx, sys, r, run, previous) found, gaveBack, err := foundAs(ctx, sys, r, raw, previous)
if err != nil { if err != nil {
return out, err return out, err
} }
@@ -1107,6 +1297,9 @@ func applyService(ctx context.Context, sys system.System, r *declaration.Service
// Read back. A service manager accepting a command says the transaction was accepted, // Read back. A service manager accepting a command says the transaction was accepted,
// not that the unit is running — one that starts and immediately dies satisfies it. // not that the unit is running — one that starts and immediately dies satisfies it.
after, err := sys.ServiceState(ctx, run, r.Unit) after, err := sys.ServiceState(ctx, run, r.Unit)
if err == nil && r.State == "running" {
after, err = stayedRunning(ctx, sys, run, r.Unit)
}
if err != nil { if err != nil {
return out, err return out, err
} }
@@ -1127,7 +1320,7 @@ func applyService(ctx context.Context, sys system.System, r *declaration.Service
} }
// Read back, for the same reason as above: a unit that starts and immediately dies // Read back, for the same reason as above: a unit that starts and immediately dies
// satisfies a service manager and nothing else. // satisfies a service manager and nothing else.
after, err := sys.ServiceState(ctx, run, r.Unit) after, err := stayedRunning(ctx, sys, run, r.Unit)
if err != nil { if err != nil {
return out, err return out, err
} }
@@ -1174,6 +1367,7 @@ func applyService(ctx context.Context, sys system.System, r *declaration.Service
func reflectOnly(ctx context.Context, sys system.System, r *declaration.Service, run Runner, func reflectOnly(ctx context.Context, sys system.System, r *declaration.Service, run Runner,
changed map[string]bool) (Outcome, error) { changed map[string]bool) (Outcome, error) {
out := begin(r) out := begin(r)
out.scope, out.user = r.Scope, r.User
out.stateless = true out.stateless = true
restart := reflected(r, changed) restart := reflected(r, changed)
reload := restartedBy(r.ReloadOn, changed) reload := restartedBy(r.ReloadOn, changed)
@@ -1217,7 +1411,7 @@ func reflectOnly(ctx context.Context, sys system.System, r *declaration.Service,
} }
// Read back: it was running, and a restart or reload that left it otherwise is a failure — // Read back: it was running, and a restart or reload that left it otherwise is a failure —
// the machine's network manager down is not a change to report and move past. // the machine's network manager down is not a change to report and move past.
after, err := sys.ServiceState(ctx, run, r.Unit) after, err := stayedRunning(ctx, sys, run, r.Unit)
if err != nil { if err != nil {
return out, err return out, err
} }
@@ -1285,16 +1479,10 @@ func remove(ctx context.Context, sys system.System, a store.Applied, run Runner,
if a.Into != nil { if a.Into != nil {
return removeInto(a) return removeInto(a)
} }
if err := os.RemoveAll(a.Target); err != nil { return removeWhole(a)
return "", "", err
}
if _, err := os.Stat(a.Target); !errors.Is(err, os.ErrNotExist) {
return "", "", fmt.Errorf("%s is still there after removing it", a.Target)
}
return "removed", "no longer declared", nil
case declaration.TypeService: case declaration.TypeService:
return removeService(ctx, sys, a, run, made[a.Target]) return removeService(ctx, sys, a, run, made[unitKey(a.Scope, a.User, a.Target)])
case declaration.TypeProcess: case declaration.TypeProcess:
// The other side of the same line: a process's unit is the host's own — it wrote the unit // The other side of the same line: a process's unit is the host's own — it wrote the unit
@@ -1337,6 +1525,10 @@ func remove(ctx context.Context, sys system.System, a store.Applied, run Runner,
// would be the data loss ADR 0030 exists to prevent, on a directory the mesh never made. // would be the data loss ADR 0030 exists to prevent, on a directory the mesh never made.
return "forgotten", "an operator-owned path is never the host's to remove", nil return "forgotten", "an operator-owned path is never the host's to remove", nil
case declaration.TypeUser:
// Kept, with its shell given back when that is safe (novox/hq ADR 0176 §2, issue 228).
return removeUser(ctx, sys, a, run)
case declaration.TypeNetwork: case declaration.TypeNetwork:
// **The reason this is a shape at all** (novox/hq ADR 0029). Orphans are removed in // **The reason this is a shape at all** (novox/hq ADR 0029). Orphans are removed in
// reverse declaration order, so a network written before the containers that join it is // reverse declaration order, so a network written before the containers that join it is
@@ -1355,16 +1547,30 @@ func remove(ctx context.Context, sys system.System, a store.Applied, run Runner,
} }
return "removed", "no longer declared", nil return "removed", "no longer declared", nil
case declaration.TypeArchive:
// Exactly what it unpacked, and the directories the host made for it once they are empty
// (novox/hq issue 162).
return removeArchive(a)
default: default:
return "", "", fmt.Errorf("%w: a %q", errNoRemoval, a.Type) return "", "", fmt.Errorf("%w: a %q", errNoRemoval, a.Type)
} }
} }
// errNoRemoval is remove's answer for a kind the host has no removal for (novox/hq issue 162): an // errNoRemoval is remove's answer for a kind the host has no removal for (novox/hq issue 162; an
// archive, among others. Fatal for an orphan the declaration dropped, so an unassignment nothing can // archive has one since). Fatal for an orphan the declaration dropped, so an unassignment nothing can
// undo is never reported as done; not fatal for a former target, which was never dropped by anyone. // undo is never reported as done; not fatal for a former target, which was never dropped by anyone.
var errNoRemoval = errors.New("no way to remove") var errNoRemoval = errors.New("no way to remove")
// exited is a command that ran and exited non-zero: its words, and the exit itself.
type exited struct {
words string
exit *exec.ExitError
}
func (e *exited) Error() string { return e.words }
func (e *exited) Unwrap() error { return e.exit }
// ExecRunner runs a real command, with stdin closed and output captured. // ExecRunner runs a real command, with stdin closed and output captured.
func ExecRunner(ctx context.Context, name string, args ...string) (string, error) { func ExecRunner(ctx context.Context, name string, args ...string) (string, error) {
cmd := exec.CommandContext(ctx, name, args...) cmd := exec.CommandContext(ctx, name, args...)
@@ -1373,8 +1579,12 @@ func ExecRunner(ctx context.Context, name string, args ...string) (string, error
if err != nil { if err != nil {
var exit *exec.ExitError var exit *exec.ExitError
if errors.As(err, &exit) { if errors.As(err, &exit) {
return string(out), fmt.Errorf("%s exited %d: %s", // The words as they always were, and the exit underneath them, so a caller asks the
name, exit.ExitCode(), strings.TrimSpace(string(exit.Stderr))) // code (system.ExitCode) rather than matching text that this line is free to reword.
return string(out), &exited{
words: fmt.Sprintf("%s exited %d: %s", name, exit.ExitCode(), strings.TrimSpace(string(exit.Stderr))),
exit: exit,
}
} }
return string(out), fmt.Errorf("%s: %w", name, err) return string(out), fmt.Errorf("%s: %w", name, err)
} }
@@ -1394,6 +1604,25 @@ func applyPackage(ctx context.Context, sys system.System, r *declaration.Package
if err != nil { if err != nil {
return out, err return out, err
} }
if r.Absent {
// Declared absent (novox/hq ADR 0180): removed when it is here, left alone when it is not.
if !installed {
out.Action = "unchanged"
out.Detail = "not installed, as declared"
return out, nil
}
if err := sys.RemovePackage(ctx, run, r.Package); err != nil {
return out, fmt.Errorf("removing %s: %w", r.Package, err)
}
if still, err := sys.PackageInstalled(ctx, run, r.Package); err != nil {
return out, err
} else if still {
return out, fmt.Errorf("%s was removed without error and the package database still has it", r.Package)
}
out.Action = "removed"
out.Detail = "declared absent; its configuration is left where the package manager leaves it"
return out, nil
}
if installed { if installed {
out.Action = "unchanged" out.Action = "unchanged"
out.Detail = "already installed" out.Detail = "already installed"
@@ -1570,9 +1799,25 @@ func containerSpecReading(r *declaration.Container, declares, reads map[string]s
for _, n := range r.Networks { for _, n := range r.Networks {
b.WriteString("also-on " + n + "\n") b.WriteString("also-on " + n + "\n")
} }
// And the capabilities it was granted (ADR 0170): one gained or dropped is a different
// container, and the runtime cannot change a running one's.
for _, c := range r.Capabilities {
b.WriteString("cap " + c + "\n")
}
// And where it logs (ADR 0179): the runtime cannot move a running container's output.
if r.Logging != "" {
b.WriteString("log " + r.Logging + "\n")
}
// The cadence is part of what was declared, so a changed schedule is a changed spec — the marker // The cadence is part of what was declared, so a changed schedule is a changed spec — the marker
// moves and the install is reported "updated" and re-established. Added only when present, so no // moves and the install is reported "updated" and re-established. Added only when present, so no
// ordinary container's or run-once step's digest moves for a field it does not set. // ordinary container's or run-once step's digest moves for a field it does not set.
// And which of its module's containers it holds still while it runs (novox/hq ADR 0189), for
// the same reason: a declaration that changed the window while the machine reported no change
// would be a machine quietly holding yesterday's containers. In declared order, which is the
// order they are stopped in.
for _, id := range r.WhileStopped {
b.WriteString("while-stopped " + id + "\n")
}
if r.Schedule != "" { if r.Schedule != "" {
b.WriteString("schedule " + r.Schedule + "\n") b.WriteString("schedule " + r.Schedule + "\n")
} }
@@ -1764,6 +2009,14 @@ func applyContainer(ctx context.Context, r *declaration.Container, run Runner,
if r.Network != "" { if r.Network != "" {
args = append(args, "--network", r.Network) args = append(args, "--network", r.Network)
} }
for _, c := range r.Capabilities {
args = append(args, "--cap-add", c)
}
if r.Logging != "" {
// The journal keeps the container's name on every line (CONTAINER_NAME), which is what a
// jail matches on (novox/hq ADR 0179); `docker logs` keeps working against the journal.
args = append(args, "--log-driver", r.Logging)
}
for _, d := range r.Dns { for _, d := range r.Dns {
args = append(args, "--dns", d) args = append(args, "--dns", d)
} }
@@ -2206,6 +2459,48 @@ func removeService(ctx context.Context, sys system.System, a store.Applied, run
// the link the operator would use to do it. // the link the operator would use to do it.
return "forgotten", "recorded before the host kept what it found; left as it is", nil return "forgotten", "recorded before the host kept what it found; left as it is", nil
} }
if a.Scope == declaration.ScopeUser {
return removeUserUnit(ctx, sys, a, run, made)
}
return giveUnitBack(ctx, sys, a, run, made)
}
// errRemovalWaits is a removal that cannot happen yet and is not a failure: the record stays, the
// outcome says why, and the next apply tries again (novox/hq ADR 0177).
var errRemovalWaits = errors.New("not removed yet")
// removeUserUnit gives back a unit in an account's own manager (novox/hq ADR 0177).
//
// **Never fatal.** A removal that fails stops the apply, and its record is there to fail the same
// way on the next one — every machine wedged on one undeclared thing, which is what issues 162 and
// 228 each ended for their own shape. An account's manager is absent for an ordinary reason, the
// person logged out, so its unit would be the commonest such wedge there is. With no manager the
// unit is kept recorded and said to wait; the first apply that finds the manager running gives it
// back. An account that is gone took its manager and its units with it, and is forgotten.
func removeUserUnit(ctx context.Context, sys system.System, a store.Applied, run Runner,
made bool) (string, string, error) {
away, err := managerAway(ctx, sys, a.Scope, a.User, run)
switch {
case errors.Is(err, errNoAccount):
return "forgotten", "the account " + a.User + " is no longer on this machine, and its manager with it", nil
case err != nil:
return "", "", fmt.Errorf("%w: %v", errRemovalWaits, err)
case away != "":
return "", "", fmt.Errorf("%w: %s; given back by the first apply after it starts", errRemovalWaits, away)
}
action, detail, err := giveUnitBack(ctx, sys, a, managerFor(a.Scope, a.User, run), made)
if err != nil {
if away, _ := managerAway(ctx, sys, a.Scope, a.User, run); away != "" {
return "", "", fmt.Errorf("%w: %s, having stopped during this apply (%v)", errRemovalWaits, away, err)
}
return "", "", fmt.Errorf("%w: in %s's own manager: %v", errRemovalWaits, a.User, err)
}
return action, detail, nil
}
// giveUnitBack is removeService's giving back, in whichever manager run reaches.
func giveUnitBack(ctx context.Context, sys system.System, a store.Applied, run Runner,
made bool) (string, string, error) {
stop := made || a.Found.State == "stopped" stop := made || a.Found.State == "stopped"
disable := made || a.Found.Boot == "disabled" disable := made || a.Found.Boot == "disabled"
@@ -2294,30 +2589,43 @@ func removeService(ctx context.Context, sys system.System, a store.Applied, run
// which the mesh never starts or stops, or of another unit altogether. What was found about one // which the mesh never starts or stops, or of another unit altogether. What was found about one
// unit says nothing about another, so a service moved to a new unit gives the old one back, just as // unit says nothing about another, so a service moved to a new unit gives the old one back, just as
// if it had been undeclared, and is read afresh for the new one; gave says what that gave back. // if it had been undeclared, and is read afresh for the new one; gave says what that gave back.
//
// A unit of the same name in another manager is another unit (novox/hq ADR 0177): a service moved
// from the machine's manager to an account's, or between accounts, gives the old one back through
// the manager it was in. run reaches the machine's manager; each side is routed from it.
func foundAs(ctx context.Context, sys system.System, r *declaration.Service, run Runner, func foundAs(ctx context.Context, sys system.System, r *declaration.Service, run Runner,
previous store.Applied) (found *store.FoundUnit, gave string, err error) { previous store.Applied) (found *store.FoundUnit, gave string, err error) {
// What a record says about its manager, only where there is a record: what an unfinished apply
// found is kept by identity alone, and was read in the manager the declaration names.
sameManager := previous.ID == "" || unitKey(previous.Scope, previous.User, "") == unitKey(r.Scope, r.User, "")
if f := previous.Found; f != nil { if f := previous.Found; f != nil {
unit := f.Unit unit := f.Unit
if unit == "" { if unit == "" {
unit = previous.Target unit = previous.Target
} }
if unit == "" || unit == r.Unit { if (unit == "" || unit == r.Unit) && sameManager {
return f, "", nil return f, "", nil
} }
// Given back as if undeclared, never as the mesh's own: whether the mesh wrote the old // Given back as if undeclared, never as the mesh's own: whether the mesh wrote the old
// unit's file is known to the removal of orphans, and that file's own record goes with it. // unit's file is known to the removal of orphans, and that file's own record goes with it.
action, detail, err := removeService(ctx, sys, action, detail, err := removeService(ctx, sys,
store.Applied{Type: string(declaration.TypeService), Target: unit, Found: f}, run, false) store.Applied{Type: string(declaration.TypeService), Target: unit, Found: f,
if err != nil { Scope: previous.Scope, User: previous.User}, run, false)
switch {
case errors.Is(err, errRemovalWaits):
// The old unit's account manager is not there to give it back to; the new unit is not
// held up for it, and the giving back is said rather than silently dropped.
gave = unit + " not given back: " + err.Error()
case err != nil:
return nil, "", fmt.Errorf("giving %s back as the host found it, now that %s is declared instead: %w", return nil, "", fmt.Errorf("giving %s back as the host found it, now that %s is declared instead: %w",
unit, r.Unit, err) unit, r.Unit, err)
} case action == "restored":
if action == "restored" {
gave = unit + " " + detail gave = unit + " " + detail
} }
} else if previous.ID != "" && !previous.Stateless && previous.Target == r.Unit { } else if previous.ID != "" && !previous.Stateless && previous.Target == r.Unit && sameManager {
return nil, "", nil return nil, "", nil
} }
run = managerFor(r.Scope, r.User, run)
state, err := sys.ServiceState(ctx, run, r.Unit) state, err := sys.ServiceState(ctx, run, r.Unit)
if err != nil { if err != nil {
return nil, gave, err return nil, gave, err
@@ -2331,23 +2639,77 @@ func foundAs(ctx context.Context, sys system.System, r *declaration.Service, run
// in. Such a unit is the mesh's, whatever was found (novox/hq ADR 0118); see removeService. // in. Such a unit is the mesh's, whatever was found (novox/hq ADR 0118); see removeService.
// //
// A drop-in is not the unit's own file, and a file the host wrote over is the machine's unit with // A drop-in is not the unit's own file, and a file the host wrote over is the machine's unit with
// the mesh's text in it: its original is kept, and put back when the file's record goes. // the mesh's text in it: its original is kept, and put back when the file's record goes — unless
// the file was changed on the machine since the mesh last wrote it, or the kept copy cannot be
// read, and then it is left as it stands and the outcome says so (removeWhole, novox/hq ADR 0118).
//
// **Keyed by manager and name** (novox/hq ADR 0177): an account's unit and the machine's of the same
// name are two units, and the mesh writing one says nothing about the other. An account's manager
// loads an administrator's units from the account's own ~/.config/systemd/user, and from
// /etc/systemd/user for every account; a unit there is the mesh's for each user-scoped record of it.
func meshMadeUnits(known store.State) map[string]bool { func meshMadeUnits(known store.State) map[string]bool {
made := map[string]bool{} made := map[string]bool{}
dirs := map[string]bool{"/etc/systemd/system": true, "/run/systemd/system": true, dirs := map[string]bool{"/etc/systemd/system": true, "/run/systemd/system": true,
filepath.Clean(unitDir): true} filepath.Clean(unitDir): true}
for _, r := range known.Resources { for _, r := range known.Resources {
if r.Type != string(declaration.TypeFile) || r.Kept != "" || r.Into != nil { if !wroteWhole(r) {
continue continue
} }
path := filepath.Clean(r.Target) path := filepath.Clean(r.Target)
if dirs[filepath.Dir(path)] { if dirs[filepath.Dir(path)] {
made[filepath.Base(path)] = true made[unitKey(declaration.ScopeSystem, "", filepath.Base(path))] = true
}
}
for _, r := range known.Resources {
if r.Type != string(declaration.TypeService) || r.Scope != declaration.ScopeUser {
continue
}
for _, path := range userUnitFiles(r.User, r.Target) {
if meshWroteWhole(known, path) {
made[unitKey(r.Scope, r.User, r.Target)] = true
}
} }
} }
return made return made
} }
// wroteWhole is a file record of a file the host wrote whole where there was none.
func wroteWhole(r store.Applied) bool {
return r.Type == string(declaration.TypeFile) && r.Kept == "" && r.Into == nil
}
// meshWroteWhole is whether this host wrote the file at path whole, where there was none.
func meshWroteWhole(known store.State, path string) bool {
path = filepath.Clean(path)
for _, r := range known.Resources {
if wroteWhole(r) && filepath.Clean(r.Target) == path {
return true
}
}
return false
}
// userUnitFiles is where an account's manager loads an administrator's unit from: the one every
// account's manager reads, and the account's own. The account's is left out when its home cannot be
// read, which only makes fewer files the mesh's.
func userUnitFiles(account, unit string) []string {
paths := []string{filepath.Join("/etc/systemd/user", unit)}
if home, err := homeOf(account); err == nil && home != "" {
paths = append(paths, filepath.Join(home, ".config", "systemd", "user", unit))
}
return paths
}
// unitKey names a unit by the manager it is in and its name (novox/hq ADR 0177): the machine's
// manager, or one account's. A system unit is the same key whether its record says "system" or
// nothing, as every record before the scope existed does.
func unitKey(scope, user, unit string) string {
if scope == declaration.ScopeUser {
return "user:" + user + "/" + unit
}
return "system/" + unit
}
// moduleOf is the module a declared resource belongs to. // moduleOf is the module a declared resource belongs to.
// //
// The mesh composes a module's resource ids as `<module>.<its own id>`, and **a module's name may // The mesh composes a module's resource ids as `<module>.<its own id>`, and **a module's name may
@@ -2369,3 +2731,169 @@ func moduleOf(identity string) (string, bool) {
} }
return identity[:at], true return identity[:at], true
} }
// handOver removes what a process replaces, once the process is running and still is a moment later
// (novox/hq issue 213, ADR 0184). A replacement that is not up keeps what it replaces in place
// and recorded, and is this apply's failure: the old one answers until a later apply finds the new
// one up.
func handOver(ctx context.Context, run Runner, resource declaration.Resource,
waiting []store.Applied, removeOrphan func(store.Applied) error, report *Report,
log func(string)) *Error {
p, ok := resource.(*declaration.Process)
if !ok {
return nil
}
if why := stillUp(ctx, run, p.Name+".service"); why != "" {
for _, orphan := range waiting {
keptForReplacement(report, log, orphan,
fmt.Sprintf("kept: %s, which replaces it, is not running (%s)", p.ID, why))
}
return &Error{Resource: p.ID, Err: fmt.Errorf(
"%s was started and is not running a moment later (%s), so what it replaces was kept: %s",
p.Name, why, appliedIDs(waiting)), Done: *report}
}
for _, orphan := range waiting {
if err := removeOrphan(orphan); err != nil {
var failed *Error
if errors.As(err, &failed) {
return failed
}
return &Error{Resource: orphan.ID, Err: err, Done: *report}
}
}
return nil
}
// handoverSettle is how long a replacement must stay up before what it replaces goes. Longer than a
// service's second look (serviceSettle): this one decides whether the last thing answering is
// removed, and a process that fails on its store or its bus does so after it opened them, not in the
// first instant. A test sets it to nothing.
var handoverSettle = 10 * time.Second
// stillUp says why a process's unit is not up and staying up, or nothing when it is: active and
// running at two looks handoverSettle apart, the same main process both times, restarted by
// nothing in between. Stricter than stayedRunning, which reads "activating" as running — the state
// a crash-looping unit is in while it waits to be started again, which is exactly the replacement
// that must not be handed anything.
func stillUp(ctx context.Context, run Runner, unit string) string {
look := func() (map[string]string, string) {
out, err := run(ctx, "systemctl", "show", unit, "--property=ActiveState", "--property=SubState",
"--property=MainPID", "--property=NRestarts")
if err != nil {
return nil, err.Error()
}
got := map[string]string{}
for _, line := range strings.Split(out, "\n") {
if key, value, found := strings.Cut(strings.TrimSpace(line), "="); found {
got[key] = value
}
}
if got["ActiveState"] != "active" || got["SubState"] != "running" {
return got, fmt.Sprintf("%s/%s", got["ActiveState"], got["SubState"])
}
return got, ""
}
first, why := look()
if why != "" {
return why
}
timer := time.NewTimer(handoverSettle)
defer timer.Stop()
select {
case <-ctx.Done():
return ctx.Err().Error()
case <-timer.C:
}
second, why := look()
if why != "" {
return why
}
if first["MainPID"] != second["MainPID"] || first["NRestarts"] != second["NRestarts"] {
return fmt.Sprintf("restarted while it was watched (pid %s → %s, restarts %s → %s)",
first["MainPID"], second["MainPID"], first["NRestarts"], second["NRestarts"])
}
return ""
}
// keptForReplacement reports an orphan kept because what replaces it is not yet in its place.
func keptForReplacement(report *Report, log func(string), orphan store.Applied, detail string) {
report.Outcomes = append(report.Outcomes, Outcome{
ID: orphan.ID, Type: orphan.Type, Target: orphan.Target, Action: "kept", Detail: detail,
})
log(fmt.Sprintf(" kept %s (%s): %s", orphan.ID, orphan.Target, strings.TrimPrefix(detail, "kept: ")))
}
func appliedIDs(applied []store.Applied) string {
ids := make([]string, 0, len(applied))
for _, a := range applied {
ids = append(ids, a.ID)
}
return strings.Join(ids, ", ")
}
// managerFor routes a unit's commands to the manager it belongs to (novox/hq ADR 0177). A system
// unit's go to the machine's manager as they always did. A user-scoped unit's go to the account's
// own: `systemctl --user --machine=<account>@`, which reaches that manager from the host's own
// process without an environment to forge or a user to switch to — and which only answers while
// the account's manager runs (a login, or lingering enabled for the account). Done on the runner
// rather than in each system: every system's reading of a unit already goes through `systemctl`,
// so this is one place instead of one per system and one per method.
//
// **The host's environment is not what reaches the account** (point 4 of the review). The
// `--machine` transport does not read XDG_RUNTIME_DIR or DBUS_SESSION_BUS_ADDRESS: it asks the
// machine's manager, over the system bus, to run `systemd-stdio-bridge --user` as the account in
// a login of its own, whose runtime directory and bus are the account's. So the host, root under
// the machine's manager with neither variable set, needs only the system bus and `systemd-run` on
// its path — both of which a systemd machine has. Only the account itself would be reached through
// its own environment instead (systemctl short-cuts to the local bus when the caller is the
// account), and the host is never the account. Being root is what lets it ask: anyone else is
// refused by the machine's manager, and the refusal is the error the unit fails with.
func managerFor(scope, user string, run Runner) Runner {
if scope != declaration.ScopeUser || user == "" {
return run
}
return func(ctx context.Context, name string, args ...string) (string, error) {
if name != "systemctl" {
return run(ctx, name, args...)
}
scoped := append([]string{"--user", "--machine=" + user + "@"}, args...)
return run(ctx, name, scoped...)
}
}
// userManagers is a service manager that runs a manager per account (novox/hq ADR 0177).
type userManagers interface {
UserManagerRunning(ctx context.Context, run system.Runner, account string) (running, exists bool, err error)
}
// errNoAccount is a user-scoped unit naming an account the machine does not have.
var errNoAccount = errors.New("no such account on this machine")
// managerAway is why the manager a unit is in cannot be reached now, or "" when it can — always
// for a system unit (novox/hq ADR 0177). Asked of the machine's manager through run, never of the
// account's, which a question would start (system.UserManagerRunning says why).
//
// Refused outright: an account the machine does not have, as ADR 0177 §1 requires, and a machine
// whose service manager has no manager per account — where a user-scoped unit would otherwise be
// applied as the machine's unit of the same name.
func managerAway(ctx context.Context, sys system.System, scope, user string, run Runner) (string, error) {
if scope != declaration.ScopeUser {
return "", nil
}
um, ok := sys.(userManagers)
if !ok {
return "", fmt.Errorf("a unit in %s's own manager, and this machine's service manager (%s) has "+
"no manager per account (novox/hq ADR 0177)", user, sys.Name())
}
running, exists, err := um.UserManagerRunning(ctx, system.Runner(run), user)
switch {
case err != nil:
return "", err
case !exists:
return "", fmt.Errorf("%w: %q, whose own manager a user-scoped unit lives in (novox/hq ADR 0177)",
errNoAccount, user)
case !running:
return user + "'s own service manager is not running: the account is not logged in and does not linger", nil
}
return "", nil
}
+117 -10
View File
@@ -56,26 +56,28 @@ func applyArchive(ctx context.Context, r *declaration.Archive, previous store.Ap
out.wrote = got out.wrote = got
// Already what it should be. The digest is the whole identity of an archive, so a matching // Already what it should be. The digest is the whole identity of an archive, so a matching
// record means the unpacked tree came from these exact bytes. // record means the unpacked tree came from these exact bytes — at this path: a record of the
if previous.Wrote == got { // same bytes somewhere else says nothing about what is here.
if previous.Wrote == got && previous.Target == r.Path {
if _, err := os.Stat(r.Path); err == nil { if _, err := os.Stat(r.Path); err == nil {
owned, err := ownedBy(r.Path, r.Owner) owned, err := ownedBy(ownershipProbe(r.Path, previous.Unpacked), r.Owner)
if err == nil && owned { if err == nil && owned {
// What it unpacked is carried, or — on a record from before the host kept it — read
// from the archive now, so the record can say it from here on (novox/hq issue 162).
out.unpacked, err = stillUnpacked(body, r.Path, previous.Unpacked)
if err != nil {
return out, err
}
return out, nil return out, nil
} }
} }
} }
if err := os.MkdirAll(r.Path, 0o755); err != nil { unpacked, written, err := replaceWith(body, r.Path, r.Owner, oursFrom(previous, r.Path))
return out, err
}
written, err := unpack(body, r.Path)
if err != nil { if err != nil {
return out, err return out, err
} }
if err := ownAll(r.Path, r.Owner); err != nil { out.unpacked = &unpacked
return out, err
}
out.Action = "updated" out.Action = "updated"
if previous.Wrote == "" { if previous.Wrote == "" {
out.Action = "created" out.Action = "created"
@@ -84,6 +86,111 @@ func applyArchive(ctx context.Context, r *declaration.Archive, previous store.Ap
return out, nil return out, nil
} }
// replaceWith makes the directory exactly the archive (novox/hq issue 220), and says what it put
// there (novox/hq issue 162).
//
// **The tree on disk is the archive and nothing else — of what the mesh put there.** The digest is
// the whole identity of what is unpacked here, so a file the previous archive had and this one does
// not must go. Unpacked over the old tree, it stayed: a bundle rebuilt as one file per entrypoint
// kept the package directory of the version before, which code could still import, and a fix that
// removed a file worked on a fresh machine only. So the archive is unpacked into a fresh directory
// beside the old one, owned, and swapped in by rename. A running process keeps the files it has open,
// and the old tree is removed only once the new one is in place. A failed unpack leaves the old tree
// untouched.
//
// **What the mesh did not put there is never swapped away** (novox/hq issue 162, ADR 0030). The
// swap is for a directory that is the host's own: one it made, holding nothing but what the mesh
// put there. A directory that was there before the archive, or that something else has written
// into since, is the machine's: the archive's files are moved into it one by one, what the previous
// archive placed and this one does not is taken out, and everything else is left as it is. A file
// the archive would write over that the mesh did not put there refuses the archive before anything
// is moved — unless it already holds exactly the archive's bytes.
func replaceWith(body []byte, path, owner string, o ours) (store.Unpacked, int, error) {
parent := filepath.Dir(path)
madeParents, err := makeDirsSaying(parent, 0o755, owner)
if err != nil {
return store.Unpacked{}, 0, err
}
parents := joinParents(madeParents, o.parents)
fresh := path + ".unpacking"
replaced := path + ".replaced"
// What an interrupted earlier attempt — or removal — left beside the directory.
for _, leftover := range []string{fresh, replaced, path + ".removing"} {
if err := os.RemoveAll(leftover); err != nil {
return store.Unpacked{}, 0, err
}
}
if err := os.Mkdir(fresh, 0o755); err != nil {
return store.Unpacked{}, 0, err
}
written, err := unpack(body, fresh)
if err == nil {
err = ownAll(fresh, owner)
}
var files, dirs []string
if err == nil {
files, dirs, err = treeOf(fresh)
}
if err != nil {
os.RemoveAll(fresh)
return store.Unpacked{}, written, err
}
info, err := os.Lstat(path)
existed := err == nil
if err != nil && !os.IsNotExist(err) {
os.RemoveAll(fresh)
return store.Unpacked{}, written, err
}
if existed && !info.IsDir() {
// Swapped, it would be deleted: a file at the path is nothing an archive put there.
os.RemoveAll(fresh)
return store.Unpacked{}, written, fmt.Errorf(
"%s is there and is not a directory, and the mesh did not put it there; nothing was unpacked", path)
}
foreign := 0
if existed && !o.all {
if foreign, err = foreignIn(path, o.paths); err != nil {
os.RemoveAll(fresh)
return store.Unpacked{}, written, err
}
}
if existed && (foreign > 0 || !(o.made || o.all)) {
u, err := mergeInto(fresh, path, owner, files, dirs, o)
os.RemoveAll(fresh)
if err != nil {
return store.Unpacked{}, written, err
}
u.Parents = parents
return u, written, nil
}
hadOne := true
if err := os.Rename(path, replaced); err != nil {
if !os.IsNotExist(err) {
os.RemoveAll(fresh)
return store.Unpacked{}, written, err
}
hadOne = false
}
if err := os.Rename(fresh, path); err != nil {
if hadOne {
// Put the old tree back rather than leave nothing at the path.
os.Rename(replaced, path)
}
os.RemoveAll(fresh)
return store.Unpacked{}, written, err
}
// The directory is the host's own: it made it, now or before, and nothing else is in it.
u := store.Unpacked{Files: files, Dirs: dirs, Made: true, Parents: parents}
if hadOne {
if err := os.RemoveAll(replaced); err != nil {
return u, written, fmt.Errorf("%s is in place, and the tree it replaced could not be removed: %w", path, err)
}
}
return u, written, nil
}
func fetch(ctx context.Context, source string) ([]byte, error) { func fetch(ctx context.Context, source string) ([]byte, error) {
request, err := http.NewRequestWithContext(ctx, http.MethodGet, source, nil) request, err := http.NewRequestWithContext(ctx, http.MethodGet, source, nil)
if err != nil { if err != nil {
+261
View File
@@ -0,0 +1,261 @@
package apply
import (
"context"
"os"
"path/filepath"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// Defends novox/hq issue 162: an archive can be undeclared. What it unpacked is gone, anything that
// was in its directory beforehand is still there, and the apply that removed it applied everything
// else in the same declaration.
func exists(path string) bool {
_, err := os.Lstat(path)
return err == nil
}
func archiveDecl(t *testing.T, id, path string, files map[string]string, more string) (string, string) {
t.Helper()
body, digest := anArchive(t, files)
return `{"id":"` + id + `","type":"archive","source":"` + serving(t, body) + `","digest":"` + digest +
`","path":"` + path + `"}` + more, digest
}
func TestUnassigningAnArchiveRemovesExactlyWhatItUnpacked(t *testing.T) {
dir := t.TempDir()
target := filepath.Join(dir, "bundles", "notes", "tools")
archive, _ := archiveDecl(t, "notes.tools", target,
map[string]string{"index.js": "x", "lib/one.js": "1", "lib/two.js": "2"},
`,{"id":"notes.conf","type":"file","path":"`+dir+`/notes.conf","content":"a"}`)
_, state, err := Apply(context.Background(), archHost(t), declare(t, archive), store.State{},
store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
if rec, _ := state.Find("notes.tools"); rec.Unpacked == nil || !rec.Unpacked.Made ||
len(rec.Unpacked.Files) != 3 || len(rec.Unpacked.Parents) != 2 {
t.Fatalf("what the archive unpacked was not recorded: %+v", rec.Unpacked)
}
// Unassigned, and the same push declares something new: both happen.
d := declare(t, `{"id":"notes.conf","type":"file","path":"`+dir+`/notes.conf","content":"a"},
{"id":"other.conf","type":"file","path":"`+dir+`/other.conf","content":"b"}`)
report, state, err := Apply(context.Background(), archHost(t), d, state, store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatalf("undeclaring an archive stopped the apply: %v", err)
}
if o := outcomeOf(report, "notes.tools"); o.Action != "removed" {
t.Fatalf("the archive: %+v", o)
}
if o := outcomeOf(report, "other.conf"); o.Action != "created" {
t.Fatalf("the rest of the declaration: %+v", o)
}
if _, still := state.Find("notes.tools"); still {
t.Fatal("the archive is still on record")
}
// The directory it was unpacked into and the parents the host made to reach it, all gone.
if exists(filepath.Join(dir, "bundles")) {
t.Fatal("what the host made for the archive is still there")
}
if !exists(filepath.Join(dir, "notes.conf")) {
t.Fatal("a file the archive did not place was removed")
}
if entries, _ := os.ReadDir(dir); len(entries) != 2 {
t.Fatalf("%d entries left in the directory", len(entries))
}
}
func TestADirectoryFoundBeforeTheArchiveKeepsWhatWasInIt(t *testing.T) {
dir := t.TempDir()
target := filepath.Join(dir, "powerlevel10k")
if err := os.MkdirAll(filepath.Join(target, "lib"), 0o755); err != nil {
t.Fatal(err)
}
os.WriteFile(filepath.Join(target, "mine.zsh"), []byte("somebody's"), 0o644)
os.WriteFile(filepath.Join(target, "lib", "mine.zsh"), []byte("somebody's"), 0o644)
archive, _ := archiveDecl(t, "shell.theme", target,
map[string]string{"p10k.zsh": "theme", "lib/theme.zsh": "lib", "gitstatus/gs": "gs"}, "")
_, state, err := Apply(context.Background(), archHost(t), declare(t, archive), store.State{},
store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
// Applying over a directory that was there keeps what was in it.
for _, kept := range []string{"mine.zsh", "lib/mine.zsh"} {
if got, _ := os.ReadFile(filepath.Join(target, kept)); string(got) != "somebody's" {
t.Fatalf("%s after the apply: %q", kept, got)
}
}
if got, _ := os.ReadFile(filepath.Join(target, "lib", "theme.zsh")); string(got) != "lib" {
t.Fatalf("lib/theme.zsh is %q", got)
}
rec, _ := state.Find("shell.theme")
if rec.Unpacked == nil || rec.Unpacked.Made || strings.Join(rec.Unpacked.Dirs, ",") != "gitstatus" {
t.Fatalf("recorded as %+v", rec.Unpacked)
}
// A new version without one of its files: that file goes, nothing else does.
archive, _ = archiveDecl(t, "shell.theme", target, map[string]string{"p10k.zsh": "theme 2"}, "")
if _, state, err = Apply(context.Background(), archHost(t), declare(t, archive), state,
store.OriginDeclared, noServices, nil, nil); err != nil {
t.Fatal(err)
}
if exists(filepath.Join(target, "lib", "theme.zsh")) || exists(filepath.Join(target, "gitstatus")) {
t.Fatal("the previous version's files are still there")
}
if !exists(filepath.Join(target, "lib", "mine.zsh")) {
t.Fatal("a file the mesh did not put there went with the previous version")
}
report, _, err := Apply(context.Background(), archHost(t), declare(t, `{"id":"other","type":"file","path":"`+dir+`/other","content":"o"}`), state,
store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
if o := outcomeOf(report, "shell.theme"); o.Action != "removed" || !strings.Contains(o.Detail, "there before") {
t.Fatalf("the archive: %+v", o)
}
if exists(filepath.Join(target, "p10k.zsh")) {
t.Fatal("the archive's file is still there")
}
for _, kept := range []string{"mine.zsh", "lib/mine.zsh"} {
if got, _ := os.ReadFile(filepath.Join(target, kept)); string(got) != "somebody's" {
t.Fatalf("%s after the removal: %q", kept, got)
}
}
}
// A file the archive would write over that the mesh did not put there refuses the archive, and the
// directory is left exactly as it was.
func TestAnArchiveDoesNotWriteOverAFileItFound(t *testing.T) {
dir := t.TempDir()
target := filepath.Join(dir, "theme")
os.MkdirAll(target, 0o755)
os.WriteFile(filepath.Join(target, "p10k.zsh"), []byte("somebody's"), 0o644)
archive, _ := archiveDecl(t, "shell.theme", target, map[string]string{"p10k.zsh": "theme", "x": "x"}, "")
if _, _, err := Apply(context.Background(), archHost(t), declare(t, archive), store.State{},
store.OriginDeclared, noServices, nil, nil); err == nil || !strings.Contains(err.Error(), "did not put there") {
t.Fatalf("written over: %v", err)
}
if got, _ := os.ReadFile(filepath.Join(target, "p10k.zsh")); string(got) != "somebody's" {
t.Fatalf("p10k.zsh is %q", got)
}
if exists(filepath.Join(target, "x")) {
t.Fatal("part of a refused archive was moved in")
}
if entries, _ := os.ReadDir(dir); len(entries) != 1 {
t.Fatalf("%d entries beside the directory", len(entries))
}
}
func TestADirectoryTheHostMadeGoesOnlyWhenEmpty(t *testing.T) {
dir := t.TempDir()
target := filepath.Join(dir, "theme")
archive, _ := archiveDecl(t, "shell.theme", target, map[string]string{"p10k.zsh": "t", "lib/a.zsh": "a"}, "")
_, state, err := Apply(context.Background(), archHost(t), declare(t, archive), store.State{},
store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
// Somebody writes into the directory the host made.
os.WriteFile(filepath.Join(target, "lib", "local.zsh"), []byte("mine"), 0o644)
report, _, err := Apply(context.Background(), archHost(t), declare(t, `{"id":"other","type":"file","path":"`+dir+`/other","content":"o"}`), state,
store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
o := outcomeOf(report, "shell.theme")
if o.Action != "removed" || !strings.Contains(o.Detail, "did not put there") {
t.Fatalf("the archive: %+v", o)
}
if exists(filepath.Join(target, "p10k.zsh")) || exists(filepath.Join(target, "lib", "a.zsh")) {
t.Fatal("the archive's files are still there")
}
if got, _ := os.ReadFile(filepath.Join(target, "lib", "local.zsh")); string(got) != "mine" {
t.Fatalf("a file the archive did not place: %q", got)
}
}
// An archive recorded before the host kept what it unpacked cannot be told from anything else in
// its directory: it is left in place, said and forgotten, and the apply goes on.
func TestAnArchiveRecordedBeforeItsFilesWereKeptIsLeftAndForgotten(t *testing.T) {
dir := t.TempDir()
target := filepath.Join(dir, "tools")
os.MkdirAll(target, 0o755)
os.WriteFile(filepath.Join(target, "index.js"), []byte("x"), 0o644)
known := store.State{Resources: []store.Applied{
{ID: "notes.tools", Type: "archive", Target: target, Wrote: "sha256:old", Origin: store.OriginDeclared},
}}
d := declare(t, `{"id":"notes.conf","type":"file","path":"`+dir+`/notes.conf","content":"a"}`)
report, state, err := Apply(context.Background(), archHost(t), d, known, store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatalf("an archive from before stopped the apply: %v", err)
}
if o := outcomeOf(report, "notes.tools"); o.Action != "forgotten" || !strings.Contains(o.Detail, "left in place") {
t.Fatalf("the archive: %+v", o)
}
if o := outcomeOf(report, "notes.conf"); o.Action != "created" {
t.Fatalf("the rest of the declaration: %+v", o)
}
if _, still := state.Find("notes.tools"); still {
t.Fatal("still on record")
}
if !exists(filepath.Join(target, "index.js")) {
t.Fatal("removed without a record of what it unpacked")
}
}
// One recorded before, still declared and unchanged, is read from its own bytes on the next apply,
// so it can be undeclared from then on: the whole directory when it holds exactly the archive, only
// the archive's files when it holds anything else.
func TestAnArchiveRecordedBeforeLearnsWhatItUnpacked(t *testing.T) {
for _, extra := range []bool{false, true} {
dir := t.TempDir()
target := filepath.Join(dir, "tools")
files := map[string]string{"index.js": "x", "lib/a.js": "a"}
archive, digest := archiveDecl(t, "notes.tools", target, files, "")
for name, body := range files {
os.MkdirAll(filepath.Dir(filepath.Join(target, name)), 0o755)
os.WriteFile(filepath.Join(target, name), []byte(body), 0o644)
}
if extra {
os.WriteFile(filepath.Join(target, "lib", "local.js"), []byte("mine"), 0o644)
}
known := store.State{Resources: []store.Applied{
{ID: "notes.tools", Type: "archive", Target: target, Wrote: digest, Origin: store.OriginDeclared},
}}
report, state, err := Apply(context.Background(), archHost(t), declare(t, archive), known,
store.OriginDeclared, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
if o := outcomeOf(report, "notes.tools"); o.Action != "unchanged" {
t.Fatalf("an unchanged archive was %s", o.Action)
}
rec, _ := state.Find("notes.tools")
if rec.Unpacked == nil || len(rec.Unpacked.Files) != 2 || rec.Unpacked.Made == extra {
t.Fatalf("extra=%v: learned %+v", extra, rec.Unpacked)
}
if _, _, err := Apply(context.Background(), archHost(t), declare(t, `{"id":"other","type":"file","path":"`+dir+`/other","content":"o"}`), state,
store.OriginDeclared, noServices, nil, nil); err != nil {
t.Fatal(err)
}
if exists(filepath.Join(target, "index.js")) || exists(filepath.Join(target, "lib", "a.js")) {
t.Fatalf("extra=%v: the archive's files are still there", extra)
}
if exists(target) != extra {
t.Fatalf("extra=%v: the directory is there: %v", extra, exists(target))
}
if extra && !exists(filepath.Join(target, "lib", "local.js")) {
t.Fatal("a file the archive did not place was removed")
}
}
}
+68
View File
@@ -0,0 +1,68 @@
package apply
import (
"context"
"os"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// novox/hq issue 220: the tree on disk is exactly the archive. A file the previous archive had and
// this one does not is gone, and nothing is left beside the directory.
func TestAnArchiveReplacesTheTreeItWasUnpackedOver(t *testing.T) {
dir := t.TempDir()
first, firstDigest := anArchive(t, map[string]string{"index.js": "old", "node_modules/dep/index.js": "dep"})
d := declare(t, `{"id":"bundle-tools","type":"archive","source":"`+serving(t, first)+
`","digest":"`+firstDigest+`","path":"`+dir+`/tools"}`)
_, state, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
second, secondDigest := anArchive(t, map[string]string{"index.js": "one file"})
d = declare(t, `{"id":"bundle-tools","type":"archive","source":"`+serving(t, second)+
`","digest":"`+secondDigest+`","path":"`+dir+`/tools"}`)
if _, _, err := Apply(context.Background(), archHost(t), d, state, store.OriginCarried, noServices, nil, nil); err != nil {
t.Fatal(err)
}
if got, _ := os.ReadFile(dir + "/tools/index.js"); string(got) != "one file" {
t.Fatalf("index.js is %q", got)
}
if _, err := os.Stat(dir + "/tools/node_modules"); err == nil {
t.Fatal("the previous archive's directory is still there")
}
entries, _ := os.ReadDir(dir)
if len(entries) != 1 {
names := []string{}
for _, e := range entries {
names = append(names, e.Name())
}
t.Fatalf("beside the tree: %v", names)
}
}
// A tree is replaced only by one that unpacked whole: an archive refused halfway leaves the old
// tree as it was.
func TestARefusedArchiveLeavesTheTreeItWouldHaveReplaced(t *testing.T) {
dir := t.TempDir()
first, firstDigest := anArchive(t, map[string]string{"index.js": "old"})
d := declare(t, `{"id":"bundle-tools","type":"archive","source":"`+serving(t, first)+
`","digest":"`+firstDigest+`","path":"`+dir+`/tools"}`)
_, state, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, noServices, nil, nil)
if err != nil {
t.Fatal(err)
}
bad, badDigest := anArchive(t, map[string]string{"index.js": "new", "../../escaped": "no"})
d = declare(t, `{"id":"bundle-tools","type":"archive","source":"`+serving(t, bad)+
`","digest":"`+badDigest+`","path":"`+dir+`/tools"}`)
if _, _, err := Apply(context.Background(), archHost(t), d, state, store.OriginCarried, noServices, nil, nil); err == nil {
t.Fatal("an escaping archive was accepted")
}
if got, _ := os.ReadFile(dir + "/tools/index.js"); string(got) != "old" {
t.Fatalf("index.js is %q after a refused archive", got)
}
if entries, _ := os.ReadDir(dir); len(entries) != 1 {
t.Fatalf("%d entries beside the tree after a refused archive", len(entries))
}
}
+506
View File
@@ -0,0 +1,506 @@
package apply
import (
"archive/tar"
"bytes"
"compress/gzip"
"crypto/sha256"
"encoding/hex"
"fmt"
"io"
"io/fs"
"os"
"path/filepath"
"sort"
"strings"
"github.com/novox/mesh-host/internal/store"
)
// What an archive put on the machine, and taking exactly that away (novox/hq issue 162).
//
// An archive unpacks many files into a directory the mesh did not necessarily make, so undeclaring
// one has a real question in it: remove what the archive put there, or remove the directory? The
// second deletes whatever else lives there — for the host's own versions directory, every other
// delivered version. So the host records what each archive unpacked and whether it made the
// directory, and removal takes away exactly that: the files the archive placed, then the
// directories the host made for them once they are empty. Never a file the archive did not place,
// never a directory that was there before, never one that still holds anything else. It is the
// rule every other kind follows: the mesh gives back what it found (ADR 0118), and data outlives
// the mesh that declared it (ADR 0030).
// ours is what of the tree at an archive's path the record says is the mesh's.
type ours struct {
// all is a record from before the host kept what an archive unpacked: since the swap of issue
// 220 the tree at the path was the archive and nothing else, so the whole of it is taken for
// the mesh's, as the swap that follows has always taken it.
all bool
// paths are the files and directories the previous archive put there, relative to the path.
paths map[string]bool
files []string
dirs []string
made bool
// parents are the directories above the path the host made to reach it, deepest first.
parents []string
}
// oursFrom reads the record of the archive before this apply, for this path only: a record of the
// same archive at a path it has moved from says nothing about what is at the new one.
func oursFrom(previous store.Applied, path string) ours {
if previous.Wrote == "" || previous.Target != path {
return ours{}
}
u := previous.Unpacked
if u == nil {
return ours{all: true, made: true}
}
o := ours{paths: map[string]bool{}, files: u.Files, dirs: u.Dirs, made: u.Made, parents: u.Parents}
for _, rel := range append(append([]string{}, u.Files...), u.Dirs...) {
o.paths[rel] = true
}
return o
}
// ownershipProbe is what says whether an archive is still its owner's. The directory, when the host
// made it; one of the archive's own files when the directory was there before — that one is held as
// found (ADR 0182), so its owner is never the archive's to judge.
func ownershipProbe(path string, u *store.Unpacked) string {
if u != nil && !u.Made && len(u.Files) > 0 {
return filepath.Join(path, u.Files[0])
}
return path
}
// stillUnpacked is what an unchanged archive has on the machine. The record's, when it has one; on
// a record from before the host kept it, read from the archive's own bytes now — and the directory
// is taken for the host's only when it holds exactly the archive and nothing else, which is what the
// swap of issue 220 leaves. Otherwise the directory is kept for somebody's, and only the archive's
// files are recorded as its.
func stillUnpacked(body []byte, path string, recorded *store.Unpacked) (*store.Unpacked, error) {
if recorded != nil {
kept := *recorded
return &kept, nil
}
files, dirs, err := listArchive(body)
if err != nil {
return nil, err
}
u := &store.Unpacked{Files: pathsOf(files)}
if exactly, err := holdsExactly(path, files, dirs); err == nil && exactly {
u.Dirs = pathsOf(dirs)
u.Made = true
}
return u, nil
}
// listArchive reads what an archive holds without unpacking it: each file's digest by its path, and
// every directory, named or implied, relative to where it unpacks. Refused on the same terms as
// unpack, so a listing never names a path an unpack would not write.
func listArchive(body []byte) (map[string]string, map[string]bool, error) {
zipped, err := gzip.NewReader(bytes.NewReader(body))
if err != nil {
return nil, nil, fmt.Errorf("this is not a gzipped tar: %w", err)
}
defer zipped.Close()
files, dirs := map[string]string{}, map[string]bool{}
reader := tar.NewReader(zipped)
for {
header, err := reader.Next()
if err == io.EOF {
return files, dirs, nil
}
if err != nil {
return nil, nil, err
}
rel := filepath.Clean(header.Name)
if rel == "." {
continue
}
if !insideRel(rel) {
return nil, nil, fmt.Errorf("%s names a path outside the archive", header.Name)
}
for d := filepath.Dir(rel); d != "."; d = filepath.Dir(d) {
dirs[filepath.ToSlash(d)] = true
}
switch header.Typeflag {
case tar.TypeDir:
dirs[filepath.ToSlash(rel)] = true
case tar.TypeReg:
sum := sha256.New()
if _, err := io.Copy(sum, io.LimitReader(reader, maxArchive)); err != nil {
return nil, nil, err
}
files[filepath.ToSlash(rel)] = hex.EncodeToString(sum.Sum(nil))
default:
return nil, nil, fmt.Errorf("%s is a %c, and this host unpacks only files and directories",
header.Name, header.Typeflag)
}
}
}
// holdsExactly is whether a directory holds the archive's files with the archive's bytes, its
// directories, and nothing else.
func holdsExactly(root string, files map[string]string, dirs map[string]bool) (bool, error) {
seen := 0
exact := true
err := filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error {
if err != nil {
return err
}
rel, err := filepath.Rel(root, path)
if err != nil {
return err
}
if rel == "." {
return nil
}
rel = filepath.ToSlash(rel)
switch {
case d.IsDir():
if !dirs[rel] {
exact = false
return filepath.SkipAll
}
case d.Type().IsRegular():
want, ok := files[rel]
if !ok || digestOfFile(path) != want {
exact = false
return filepath.SkipAll
}
seen++
default:
exact = false
return filepath.SkipAll
}
return nil
})
if err != nil {
return false, err
}
return exact && seen == len(files), nil
}
func digestOfFile(path string) string {
file, err := os.Open(path)
if err != nil {
return ""
}
defer file.Close()
sum := sha256.New()
if _, err := io.Copy(sum, file); err != nil {
return ""
}
return hex.EncodeToString(sum.Sum(nil))
}
// treeOf is every file and directory under root, relative to it, slash-separated and sorted.
func treeOf(root string) (files, dirs []string, err error) {
err = filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error {
if err != nil {
return err
}
rel, err := filepath.Rel(root, path)
if err != nil || rel == "." {
return err
}
if d.IsDir() {
dirs = append(dirs, filepath.ToSlash(rel))
} else {
files = append(files, filepath.ToSlash(rel))
}
return nil
})
sort.Strings(files)
sort.Strings(dirs)
if files == nil {
files = []string{}
}
return files, dirs, err
}
// foreignIn counts what under root the mesh did not put there. A directory that is not the mesh's
// counts once, with everything in it.
func foreignIn(root string, mine map[string]bool) (int, error) {
count := 0
err := filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error {
if err != nil {
return err
}
rel, err := filepath.Rel(root, path)
if err != nil || rel == "." {
return err
}
if !mine[filepath.ToSlash(rel)] {
count++
if d.IsDir() {
return filepath.SkipDir
}
}
return nil
})
return count, err
}
// mergeInto moves a freshly unpacked archive into a directory that is not only the mesh's, one entry
// at a time, and takes out what the previous archive placed that this one does not. Everything the
// mesh did not put there stays as it is. Collisions are looked for before anything is moved, so a
// refused archive leaves the directory exactly as it was.
func mergeInto(fresh, path, owner string, files, dirs []string, o ours) (store.Unpacked, error) {
var collisions []string
for _, rel := range dirs {
info, err := os.Lstat(filepath.Join(path, filepath.FromSlash(rel)))
if err == nil && !info.IsDir() && !o.paths[rel] {
collisions = append(collisions, rel)
}
}
for _, rel := range files {
dest := filepath.Join(path, filepath.FromSlash(rel))
info, err := os.Lstat(dest)
switch {
case err != nil:
case o.paths[rel] && !info.IsDir():
case info.IsDir():
if !o.paths[rel] {
collisions = append(collisions, rel)
} else if n, err := foreignIn(dest, o.paths); err != nil || n > 0 {
collisions = append(collisions, rel)
}
case !info.Mode().IsRegular() ||
digestOfFile(dest) != digestOfFile(filepath.Join(fresh, filepath.FromSlash(rel))):
// Already holding exactly the archive's bytes is not a collision: it is what an
// interrupted earlier apply of this same archive left, or the same file either way.
collisions = append(collisions, rel)
}
}
if len(collisions) > 0 {
shown := collisions
if len(shown) > 5 {
shown = shown[:5]
}
return store.Unpacked{}, fmt.Errorf("%s already holds %d path(s) the archive would write over "+
"and the mesh did not put there (%s); nothing was unpacked, and what is there is left as it "+
"is (novox/hq issue 162)", path, len(collisions), strings.Join(shown, ", "))
}
var made []string
for _, rel := range dirs {
dest := filepath.Join(path, filepath.FromSlash(rel))
info, err := os.Lstat(dest)
if err == nil && info.IsDir() {
if o.paths[rel] {
made = append(made, rel)
}
continue
}
if err == nil {
// The previous archive's file where this one has a directory.
if err := os.Remove(dest); err != nil {
return store.Unpacked{}, err
}
}
mode := os.FileMode(0o755)
if from, err := os.Stat(filepath.Join(fresh, filepath.FromSlash(rel))); err == nil {
mode = from.Mode().Perm()
}
if err := os.Mkdir(dest, mode); err != nil {
return store.Unpacked{}, err
}
if err := own(dest, owner); err != nil {
return store.Unpacked{}, err
}
made = append(made, rel)
}
for _, rel := range files {
dest := filepath.Join(path, filepath.FromSlash(rel))
if info, err := os.Lstat(dest); err == nil && info.IsDir() {
// The previous archive's directory where this one has a file, holding nothing else.
if err := os.RemoveAll(dest); err != nil {
return store.Unpacked{}, err
}
}
if err := os.Rename(filepath.Join(fresh, filepath.FromSlash(rel)), dest); err != nil {
return store.Unpacked{}, err
}
}
// What the previous archive placed and this one does not.
now := map[string]bool{}
for _, rel := range append(append([]string{}, files...), dirs...) {
now[rel] = true
}
for _, rel := range o.files {
if now[rel] {
continue
}
if err := os.Remove(filepath.Join(path, filepath.FromSlash(rel))); err != nil && !os.IsNotExist(err) {
return store.Unpacked{}, err
}
}
for _, rel := range deepestFirst(o.dirs) {
if now[rel] {
continue
}
dest := filepath.Join(path, filepath.FromSlash(rel))
if err := os.Remove(dest); err != nil && !os.IsNotExist(err) {
// Still holding something the mesh did not put there: kept, and still the host's to
// take away once it is empty.
made = append(made, rel)
}
}
sort.Strings(made)
return store.Unpacked{Files: files, Dirs: made, Made: o.made}, nil
}
// removeArchive is what undeclaring an archive does (novox/hq issue 162): exactly the files it
// unpacked, then the directories the host made for them once they are empty.
//
// **Never fatal.** An archive that could not be removed stopped the whole apply, on every apply
// after, until it was declared again — so every module with tools was un-unassignable, and a race
// between two pushes froze a machine against every other change. Whatever cannot be taken away is
// said, left in place, and forgotten, as a former target is (issue 194).
//
// **Removed whole when it is the host's own, and in one step.** A directory the host made that
// holds nothing but the archive is renamed aside and then removed: a reader — the runtime serving a
// module's tools from its bundle — sees the whole tree or none of it, never half, and a file it has
// open stays readable until it closes it. The runtime is told the module went by its own membership,
// not by the files disappearing.
func removeArchive(a store.Applied) (string, string, error) {
if store.IsFormer(a.ID) {
// The version before is what a rollback starts (ADR 0141) and what a reader may still have
// open; the launcher retires the host's own versions, not the apply (issue 194).
return "forgotten", "a former target left in place: only an archive the declaration dropped is " +
"taken away (novox/hq issues 162, 194)", nil
}
u := a.Unpacked
if u == nil {
return "forgotten", "left in place: recorded before the host kept what an archive unpacked, so " +
"its files cannot be told from anything else there (novox/hq issue 162)", nil
}
root := filepath.Clean(a.Target)
mine := map[string]bool{}
for _, rel := range append(append([]string{}, u.Files...), u.Dirs...) {
if !insideRel(filepath.FromSlash(rel)) {
return "forgotten", fmt.Sprintf("left in place: its record names %q, which is not inside %s",
rel, root), nil
}
mine[rel] = true
}
info, err := os.Lstat(root)
if os.IsNotExist(err) {
removeParents(root, u.Parents)
return "forgotten", "no longer there", nil
}
if err != nil {
return "forgotten", fmt.Sprintf("left in place: %v", err), nil
}
if !info.IsDir() {
return "forgotten", "left in place: no longer a directory, so not what the archive was unpacked into", nil
}
foreign, err := foreignIn(root, mine)
if err != nil {
return "forgotten", fmt.Sprintf("left in place: cannot read what is in it: %v", err), nil
}
if u.Made && foreign == 0 {
aside := root + ".removing"
if err := os.RemoveAll(aside); err == nil {
if err := os.Rename(root, aside); err == nil {
if err := os.RemoveAll(aside); err != nil {
return "forgotten", fmt.Sprintf("taken out of place, and what it unpacked could not be "+
"removed from %s: %v — remove it by hand", aside, err), nil
}
removeParents(root, u.Parents)
return "removed", fmt.Sprintf("no longer declared; the %d file(s) it unpacked, and the "+
"directory the host made for them", len(u.Files)), nil
}
}
// A rename that could not be made is taken file by file instead.
}
removed := 0
var failed []string
for _, rel := range u.Files {
err := os.Remove(filepath.Join(root, filepath.FromSlash(rel)))
switch {
case err == nil:
removed++
case os.IsNotExist(err):
default:
failed = append(failed, err.Error())
}
}
for _, rel := range deepestFirst(u.Dirs) {
// Only once empty: what is still inside is somebody's.
_ = os.Remove(filepath.Join(root, filepath.FromSlash(rel)))
}
detail := fmt.Sprintf("no longer declared; %d file(s) it unpacked removed", removed)
switch {
case u.Made && os.Remove(root) == nil:
removeParents(root, u.Parents)
detail += ", and the directory the host made for them"
case u.Made:
left, _ := os.ReadDir(root)
detail += fmt.Sprintf("; the directory is kept: %d item(s) inside that the mesh did not put there",
len(left))
default:
detail += "; the directory is kept: it was there before the archive"
}
if len(failed) > 0 {
// Said and not fatal: fatal, the record would stay and fail the same way on every apply
// after — the very wedge this removal exists to end.
return "forgotten", detail + "; could not remove, and left in place: " + strings.Join(failed, "; "), nil
}
return "removed", detail, nil
}
// removeParents takes away the directories above an archive the host made to reach it, deepest
// first, each only once it is empty and only if it is above the archive's directory.
func removeParents(root string, parents []string) {
for _, p := range parents {
clean := filepath.Clean(p)
if !filepath.IsAbs(clean) || !strings.HasPrefix(root, clean+string(os.PathSeparator)) {
continue
}
_ = os.Remove(clean)
}
}
// joinParents is the parents made now and those recorded before, deepest first, once each.
func joinParents(now, before []string) []string {
seen := map[string]bool{}
var out []string
for _, p := range append(append([]string{}, now...), before...) {
if !seen[p] {
seen[p] = true
out = append(out, p)
}
}
sort.Slice(out, func(i, j int) bool { return len(out[i]) > len(out[j]) })
return out
}
// insideRel is whether a relative path stays inside the directory it is relative to.
func insideRel(rel string) bool {
clean := filepath.Clean(rel)
return clean != "." && !filepath.IsAbs(clean) && clean != ".." &&
!strings.HasPrefix(clean, ".."+string(os.PathSeparator))
}
func deepestFirst(rels []string) []string {
out := append([]string{}, rels...)
sort.Slice(out, func(i, j int) bool {
return strings.Count(out[i], "/") > strings.Count(out[j], "/") ||
(strings.Count(out[i], "/") == strings.Count(out[j], "/") && out[i] > out[j])
})
return out
}
func pathsOf[V any](m map[string]V) []string {
out := make([]string, 0, len(m))
for k := range m {
out = append(out, k)
}
sort.Strings(out)
return out
}
+1 -1
View File
@@ -183,7 +183,7 @@ func applyBlock(r *declaration.File, previous store.Applied) (Outcome, error) {
} else if mode, err = modeOf(r.Mode, mode); err != nil { } else if mode, err = modeOf(r.Mode, mode); err != nil {
return out, err return out, err
} }
if err := os.MkdirAll(filepath.Dir(real), 0o755); err != nil { if err := makeDirs(filepath.Dir(real), 0o755, r.Owner); err != nil {
return out, err return out, err
} }
if err := writeAtomically(real, []byte(next), mode); err != nil { if err := writeAtomically(real, []byte(next), mode); err != nil {
+6 -13
View File
@@ -10,9 +10,9 @@ import (
// The host's own former archive stops nothing (novox/hq issue 194). A new host's first apply finds // The host's own former archive stops nothing (novox/hq issue 194). A new host's first apply finds
// the version before it as a former target of the archive that delivered it; an archive has no // the version before it as a former target of the archive that delivered it; an archive has no
// removal (issue 162), and the refusal stopped every machine applying anything. A former target of // removal then (issue 162), and the refusal stopped every machine applying anything. A former
// such a kind is left in place, said, and forgotten. An archive the declaration dropped still fails, // target of an archive is left in place, said, and forgotten: the version before is what a rollback
// as 162 has it. // starts (ADR 0141).
func TestTheHostsOwnFormerArchiveIsLeftInPlaceNotFatal(t *testing.T) { func TestTheHostsOwnFormerArchiveIsLeftInPlaceNotFatal(t *testing.T) {
run := func(_ context.Context, name string, args ...string) (string, error) { run := func(_ context.Context, name string, args ...string) (string, error) {
if name == "docker" && args[0] == "info" { if name == "docker" && args[0] == "info" {
@@ -62,14 +62,7 @@ func TestTheHostsOwnFormerArchiveIsLeftInPlaceNotFatal(t *testing.T) {
t.Fatalf("the rest of the declaration was not applied: %+v", report.Outcomes) t.Fatalf("the rest of the declaration was not applied: %+v", report.Outcomes)
} }
// An archive the declaration dropped is a different matter: nothing can undo it, and saying // An archive the declaration dropped is taken away since issue 162; one recorded before the
// it was would report an effect the host declined to have (issue 162). // host kept what it unpacked is left in place and forgotten, never fatal — that case is
dropped := store.State{Resources: []store.Applied{ // archive_removal_test.go's.
{ID: "tool.next", Type: "archive", Target: "/usr/lib/tool/versions/old", Origin: store.OriginDeclared},
}}
only := parse(t, `{"declaration":1,"resources":[{"id":"notes.conf","type":"file","path":"`+dir+`/notes.conf","content":"x"}]}`)
if _, _, err := Apply(context.Background(), archHost(t), only, dropped, store.OriginDeclared, run, nil, nil); err == nil ||
!strings.Contains(err.Error(), "no way to remove") {
t.Fatalf("a dropped archive was passed over: %v", err)
}
} }
+2 -2
View File
@@ -122,8 +122,8 @@ func TestOnlyAUnitsOwnFileTheMeshCreatedMakesItTheMeshs(t *testing.T) {
made := meshMadeUnits(known) made := meshMadeUnits(known)
for unit, want := range map[string]bool{"made.service": true, "runtime.service": true, "kept.service": false, for unit, want := range map[string]bool{"made.service": true, "runtime.service": true, "kept.service": false,
"into.service": false, "mesh.conf": false, "docker.service.d": false, "elsewhere.service": false, "dir.service": false} { "into.service": false, "mesh.conf": false, "docker.service.d": false, "elsewhere.service": false, "dir.service": false} {
if made[unit] != want { if made[unitKey("", "", unit)] != want {
t.Errorf("%s: made %v, want %v", unit, made[unit], want) t.Errorf("%s: made %v, want %v", unit, made[unitKey("", "", unit)], want)
} }
} }
} }
+318
View File
@@ -0,0 +1,318 @@
package apply
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// novox/hq issue 213: the controller moves from a container to a process on the one machine that
// runs it. Every orphan is removed before anything is applied, so without a handover the container
// went first and nothing answered the mesh's verbs while the process was fetched, unpacked and
// started — and for ever, if it did not start. A process that `replaces` the container is applied
// first; the container goes only once the process is running a moment later.
// aMachine fakes the service manager and the container runtime: it records every command, answers
// `systemctl show` with whether the process is running, and has a container until it is removed.
type aMachine struct {
commands []string
running bool // what `systemctl show` says of the process once it was started
started bool
container bool
timer bool // whether a timer, once started, is up
crashing bool // up at the first look, waiting to restart at the next
looks int
}
func (m *aMachine) run(ctx context.Context, name string, args ...string) (string, error) {
line := name + " " + strings.Join(args, " ")
m.commands = append(m.commands, line)
switch {
case name == "systemctl" && len(args) > 0 && (args[0] == "restart" || args[0] == "start"):
m.started = true
case name == "systemctl" && len(args) > 0 && args[0] == "is-active" && strings.HasSuffix(args[len(args)-1], ".timer"):
if m.timer {
return "active", nil
}
return "inactive", errors.New("inactive")
case name == "systemctl" && len(args) > 0 && args[0] == "is-active":
if m.started && m.running {
return "active", nil
}
return "inactive", errors.New("inactive")
case name == "systemctl" && len(args) > 0 && args[0] == "show":
m.looks++
if m.crashing && m.started {
// Up at the first look; waiting to be started again, a new process, at the second.
if m.looks == 1 {
return "ActiveState=active\nSubState=running\nMainPID=42\nNRestarts=0\n", nil
}
return "ActiveState=activating\nSubState=auto-restart\nMainPID=0\nNRestarts=1\n", nil
}
if m.started && m.running {
return "ActiveState=active\nSubState=running\nMainPID=42\nNRestarts=0\n", nil
}
return "ActiveState=inactive\nSubState=dead\nMainPID=0\nNRestarts=0\n", nil
case name == "docker" && len(args) > 1 && args[0] == "rm":
m.container = false
case name == "docker" && len(args) > 1 && args[0] == "container" && args[1] == "inspect":
if !m.container {
return "", errors.New("no such container")
}
return "true\t", nil
}
return "", nil
}
func (m *aMachine) index(prefix string) int {
for i, c := range m.commands {
if strings.HasPrefix(c, prefix) {
return i
}
}
return -1
}
// theController is a machine whose controller ran as a container, recorded, and a declaration that
// runs it as a process from a bundle served here instead.
func theController(t *testing.T, digest, source, extra string) (store.State, string) {
t.Helper()
known := store.State{Resources: []store.Applied{{
Origin: store.OriginDeclared, ID: "mesh-controller.server", Type: "container", Target: "mesh-controller",
}}}
return known, `{"declaration":1,"resources":[
{"id":"mesh-controller.controller","type":"process","name":"mesh-controller","source":"` + source +
`","digest":"` + digest + `","run":["./mesh-controller","serve"]` + extra + `}]}`
}
func onAMachine(t *testing.T) {
t.Helper()
serviceSettle, handoverSettle = 0, 0
wasUnits, wasBundles := unitDir, daemonRoot
unitDir, daemonRoot = t.TempDir(), t.TempDir()
t.Cleanup(func() { unitDir, daemonRoot = wasUnits, wasBundles })
}
func TestAContainerAProcessReplacesGoesOnlyOnceTheProcessRuns(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
known, raw := theController(t, digest, serving(t, body), `,"replaces":["mesh-controller.server"]`)
m := &aMachine{running: true, container: true}
report, after, err := Apply(context.Background(), archHost(t), parse(t, raw), known,
store.OriginDeclared, m.run, nil, nil)
if err != nil {
t.Fatalf("the handover failed: %v", err)
}
started, removed := m.index("systemctl restart mesh-controller.service"), m.index("docker rm -f mesh-controller")
if started < 0 || removed < 0 {
t.Fatalf("the process was not started or the container not removed: %v", m.commands)
}
if removed < started {
t.Fatalf("the container was removed before its replacement was started — a window with "+
"nothing answering: %v", m.commands)
}
if looked := m.index("systemctl show mesh-controller.service"); looked < 0 || looked > removed {
t.Errorf("the container was removed without looking whether the process runs: %v", m.commands)
}
if o := outcomeOf(report, "mesh-controller.server"); o.Action != "removed" {
t.Errorf("the container's outcome is %+v, want removed", o)
}
if _, still := after.Find("mesh-controller.server"); still {
t.Error("the host still records the container it removed")
}
if _, has := after.Find("mesh-controller.controller"); !has {
t.Error("the process was not recorded")
}
}
// The case the handover exists for: the replacement does not stay up. The container keeps
// answering, stays recorded so a later apply hands it over, and the apply says why it failed.
func TestAContainerIsKeptWhenItsReplacementDoesNotRun(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
known, raw := theController(t, digest, serving(t, body), `,"replaces":["mesh-controller.server"]`)
m := &aMachine{running: false, container: true}
report, after, err := Apply(context.Background(), archHost(t), parse(t, raw), known,
store.OriginDeclared, m.run, nil, nil)
if err == nil {
t.Fatal("a replacement that is not running was reported as a clean apply")
}
if !strings.Contains(err.Error(), "mesh-controller.server") {
t.Errorf("the failure does not name what was kept: %v", err)
}
if m.index("docker rm") >= 0 {
t.Fatalf("the container was removed though its replacement is not running: %v", m.commands)
}
if o := outcomeOf(report, "mesh-controller.server"); o.Action != "kept" {
t.Errorf("the container's outcome is %+v, want kept", o)
}
if _, still := after.Find("mesh-controller.server"); !still {
t.Fatal("the container was forgotten, so no later apply would ever remove it")
}
// The next apply finds the process up and finishes the handover.
m.running, m.commands = true, nil
report, after, err = Apply(context.Background(), archHost(t), parse(t, raw), after,
store.OriginDeclared, m.run, nil, nil)
if err != nil {
t.Fatalf("the second apply failed: %v", err)
}
if o := outcomeOf(report, "mesh-controller.server"); o.Action != "removed" {
t.Errorf("the second apply did not hand over: %+v (%v)", o, m.commands)
}
if _, still := after.Find("mesh-controller.server"); still {
t.Error("the container is still recorded after the handover")
}
}
// A replacement that never applied — its bundle is not what was declared — touches nothing, and
// what it replaces keeps running.
func TestAContainerIsKeptWhenItsReplacementFailsToApply(t *testing.T) {
onAMachine(t)
body, _ := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
known, raw := theController(t, "sha256:"+strings.Repeat("b", 64), serving(t, body),
`,"replaces":["mesh-controller.server"]`)
m := &aMachine{running: true, container: true}
report, after, err := Apply(context.Background(), archHost(t), parse(t, raw), known,
store.OriginDeclared, m.run, nil, nil)
if err == nil {
t.Fatal("a replacement whose bundle did not match was reported applied")
}
if m.index("docker rm") >= 0 {
t.Fatalf("the container was removed though nothing replaced it: %v", m.commands)
}
if o := outcomeOf(report, "mesh-controller.server"); o.Action != "kept" {
t.Errorf("the container's outcome is %+v, want kept", o)
}
if _, still := after.Find("mesh-controller.server"); !still {
t.Fatal("the container was forgotten")
}
}
// Without `replaces` nothing changes: an orphan goes before anything is applied, as it always has.
// Kept as a test because it is the window the field exists to close.
func TestWithoutReplacesAnOrphanStillGoesFirst(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
known, raw := theController(t, digest, serving(t, body), ``)
m := &aMachine{running: true, container: true}
if _, _, err := Apply(context.Background(), archHost(t), parse(t, raw), known,
store.OriginDeclared, m.run, nil, nil); err != nil {
t.Fatal(err)
}
if removed, started := m.index("docker rm -f mesh-controller"), m.index("systemctl restart mesh-controller.service"); removed < 0 || removed > started {
t.Fatalf("an orphan nothing replaces was not removed first: %v", m.commands)
}
}
// novox/hq issue 213, beyond what the oneshot unit (process_step_test.go) already holds: a step
// written `./name` runs its own bundle's binary — the controller's preparation is its own binary —
// and a step is started, never enabled.
func TestAStepRunsItsOwnBundlesBinaryAndIsNotEnabled(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
m := &aMachine{running: true}
d := declare(t, `{"id":"mesh-controller.controller-prepare","type":"process","name":"mesh-controller-prepare",
"source":"`+serving(t, body)+`","digest":"`+digest+`","run":["./mesh-controller","prepare"],"run-once":true}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared, m.run, nil, nil); err != nil {
t.Fatalf("the step failed: %v", err)
}
for _, c := range m.commands {
if strings.HasPrefix(c, "./") || strings.HasPrefix(c, "systemctl enable") {
t.Errorf("the step was run directly or enabled: %v", m.commands)
}
}
unit, err := os.ReadFile(filepath.Join(unitDir, "mesh-controller-prepare.service"))
if err != nil {
t.Fatal(err)
}
want := "ExecStart=" + filepath.Join(daemonRoot, "mesh-controller-prepare", "mesh-controller") + " prepare"
if !strings.Contains(string(unit), want) {
t.Errorf("the step does not run its own bundle's binary (%q):\n%s", want, unit)
}
}
func TestAStepThatFailsGatesItsModule(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
src := serving(t, body)
run := func(ctx context.Context, name string, args ...string) (string, error) {
if name == "systemctl" && len(args) > 1 && args[0] == "start" {
return "", errors.New("Job for mesh-controller-prepare.service failed")
}
if name == "systemctl" && len(args) > 0 && args[0] == "restart" {
t.Errorf("the module's process was started after its step failed")
}
return "", nil
}
d := declare(t, `{"id":"mesh-controller.controller-prepare","type":"process","name":"mesh-controller-prepare",
"source":"`+src+`","digest":"`+digest+`","run":["./mesh-controller","prepare"],"run-once":true},
{"id":"mesh-controller.controller","type":"process","name":"mesh-controller",
"source":"`+src+`","digest":"`+digest+`","run":["./mesh-controller","serve"]}`)
report, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared, run, nil, nil)
if err == nil {
t.Fatal("a failed step was reported as a clean apply")
}
if o := outcomeOf(report, "mesh-controller.controller"); o.Action != "skipped" {
t.Errorf("the process after a failed step was %+v, want skipped", o)
}
}
// A completed step is done: applied again unchanged, it is not run again — its service is never up
// between runs, and reading that as "a daemon that stopped" re-ran the controller's preparation on
// every apply. Nor is a scheduled run started off its cadence; its timer is what is kept up.
func TestACompletedStepIsNotRunAgainAndAScheduleIsItsTimer(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"job": "#!/bin/sh\n"})
src := serving(t, body)
for _, mode := range []string{`"run-once":true`, `"schedule":"0 3 * * *"`} {
// The service of either is never up between runs; a scheduled one's timer is.
m := &aMachine{running: false, timer: true}
d := declare(t, `{"id":"m.job","type":"process","name":"m-job","source":"`+src+`","digest":"`+digest+
`","run":["./job"],`+mode+`}`)
_, state, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared, m.run, nil, nil)
if err != nil {
t.Fatalf("%s: first apply: %v", mode, err)
}
m.commands = nil
report, _, err := Apply(context.Background(), archHost(t), d, state, store.OriginDeclared, m.run, nil, nil)
if err != nil {
t.Fatalf("%s: second apply: %v", mode, err)
}
if m.index("systemctl start m-job.service") >= 0 || m.index("systemctl restart m-job.service") >= 0 {
t.Errorf("%s: an unchanged apply ran the job again: %v", mode, m.commands)
}
if o := outcomeOf(report, "m.job"); o.Action != "unchanged" {
t.Errorf("%s: an unchanged apply reported %+v", mode, o)
}
}
}
// A replacement crash-looping between its restarts is not up, though the service manager calls it
// "activating" — the state the host's ordinary second look accepts as running. Removing the
// container on that reading would leave nothing answering.
func TestAContainerIsKeptWhenItsReplacementIsCrashLooping(t *testing.T) {
onAMachine(t)
body, digest := anArchive(t, map[string]string{"mesh-controller": "#!/bin/sh\n"})
known, raw := theController(t, digest, serving(t, body), `,"replaces":["mesh-controller.server"]`)
m := &aMachine{crashing: true, container: true}
_, after, err := Apply(context.Background(), archHost(t), parse(t, raw), known,
store.OriginDeclared, m.run, nil, nil)
if err == nil || !strings.Contains(err.Error(), "auto-restart") {
t.Fatalf("a crash-looping replacement was accepted: %v", err)
}
if m.index("docker rm") >= 0 {
t.Fatalf("the container was removed for a replacement that keeps dying: %v", m.commands)
}
if _, still := after.Find("mesh-controller.server"); !still {
t.Fatal("the container was forgotten")
}
}
+62 -9
View File
@@ -109,9 +109,30 @@ func lookBefore(ctx context.Context, sys system.System, d *declaration.Declarati
} }
} }
case *declaration.Service: case *declaration.Service:
if res.Stateless() || known.Recorded(string(declaration.TypeService), res.Unit) { if res.Stateless() || recordedUnit(known, res) {
continue continue
} }
key := "unit:" + unitKey(res.Scope, res.User, res.Unit)
run := run
if res.UserScoped() {
// In the account's own manager (novox/hq ADR 0177), and only asked while it runs. With
// it not running, what can be known is whether somebody put the unit's file where that
// manager loads an administrator's units from — and the mesh's own file there is not
// somebody's. An account the machine does not have has no unit to find.
away, err := managerAway(ctx, sys, res.Scope, res.User, run)
if err != nil {
continue
}
if away != "" {
for _, path := range userUnitFiles(res.User, res.Unit) {
if present(path) && !meshWroteWhole(known, path) {
seen.is[key] = true
}
}
continue
}
run = managerFor(res.Scope, res.User, run)
}
// **Found is a unit somebody put on this machine, or one the machine uses.** // **Found is a unit somebody put on this machine, or one the machine uses.**
// //
// Where it comes from first: a unit the service manager loads from outside /usr — // Where it comes from first: a unit the service manager loads from outside /usr —
@@ -129,14 +150,14 @@ func lookBefore(ctx context.Context, sys system.System, d *declaration.Declarati
continue continue
} }
if from, ok := sys.(unitFiles); ok { if from, ok := sys.(unitFiles); ok {
if path, err := from.ServiceUnitFile(ctx, run, res.Unit); err == nil && installedByHand(path) { if path, err := from.ServiceUnitFile(ctx, run, res.Unit); err == nil && installedByHand(known, path) {
seen.is["unit:"+res.Unit] = true seen.is[key] = true
continue continue
} }
} }
boot, _ := sys.ServiceBoot(ctx, run, res.Unit) boot, _ := sys.ServiceBoot(ctx, run, res.Unit)
if state == "running" || boot == "enabled" { if state == "running" || boot == "enabled" {
seen.is["unit:"+res.Unit] = true seen.is[key] = true
} }
case *declaration.Container: case *declaration.Container:
if known.Recorded(string(declaration.TypeContainer), res.Name) { if known.Recorded(string(declaration.TypeContainer), res.Name) {
@@ -177,12 +198,26 @@ type unitFiles interface {
} }
// installedByHand is whether a unit file is one somebody put on this machine rather than one a // installedByHand is whether a unit file is one somebody put on this machine rather than one a
// package ships: anywhere but /usr, where distributions keep what they install. // package ships: anywhere but /usr, where distributions keep what they install — and not one this
func installedByHand(path string) bool { // host wrote whole. That covers an account's units (novox/hq ADR 0177): one under the account's
// ~/.config/systemd/user or /etc/systemd/user is somebody's, unless the mesh wrote it there.
func installedByHand(known store.State, path string) bool {
if path == "" { if path == "" {
return false return false
} }
return !strings.HasPrefix(filepath.Clean(path), "/usr/") return !strings.HasPrefix(filepath.Clean(path), "/usr/") && !meshWroteWhole(known, path)
}
// recordedUnit is whether this host has a record of a service in the same manager as res, under the
// same name (novox/hq ADR 0177): an account's unit and the machine's of one name are two units.
func recordedUnit(known store.State, res *declaration.Service) bool {
want := unitKey(res.Scope, res.User, res.Unit)
for _, r := range known.Resources {
if r.Type == string(declaration.TypeService) && unitKey(r.Scope, r.User, r.Target) == want {
return true
}
}
return false
} }
func present(path string) bool { func present(path string) bool {
@@ -374,7 +409,7 @@ func holdOnAdopted(ctx context.Context, sys system.System, r declaration.Resourc
case *declaration.User: case *declaration.User:
isFound = before.has("user:" + res.Name) isFound = before.has("user:" + res.Name)
case *declaration.Service: case *declaration.Service:
isFound = before.has("unit:" + res.Unit) isFound = before.has("unit:" + unitKey(res.Scope, res.User, res.Unit))
} }
} }
if !isFound { if !isFound {
@@ -388,7 +423,11 @@ func holdOnAdopted(ctx context.Context, sys system.System, r declaration.Resourc
// A held service is not started, stopped, enabled or restarted — but a reload stops nothing, // A held service is not started, stopped, enabled or restarted — but a reload stops nothing,
// so one the module names still happens (novox/hq ADR 0102, ADR 0103). // so one the module names still happens (novox/hq ADR 0102, ADR 0103).
if svc, ok := r.(*declaration.Service); ok && svc.State == "running" { if svc, ok := r.(*declaration.Service); ok && svc.State == "running" {
if which := restartedBy(svc.ReloadOn, changed); len(which) > 0 { // In the unit's own manager, and nothing to reload in one that is not running (novox/hq ADR
// 0177): the state read below then fails, and the reload is not attempted.
away, awayErr := managerAway(ctx, sys, svc.Scope, svc.User, run)
run := managerFor(svc.Scope, svc.User, run)
if which := restartedBy(svc.ReloadOn, changed); len(which) > 0 && awayErr == nil && away == "" {
if state, err := sys.ServiceState(ctx, run, svc.Unit); err == nil && state == "running" { if state, err := sys.ServiceState(ctx, run, svc.Unit); err == nil && state == "running" {
reloader, can := sys.(serviceReloader) reloader, can := sys.(serviceReloader)
if !can { if !can {
@@ -711,6 +750,20 @@ func hold(ctx context.Context, sys system.System, r declaration.Resource, module
} }
detail = "the user was found on the machine; its shell and groups are kept until " + module + " is taken" detail = "the user was found on the machine; its shell and groups are kept until " + module + " is taken"
case *declaration.Service: case *declaration.Service:
if res.UserScoped() {
// Read in the account's own manager, and not at all while it is not running (novox/hq
// ADR 0177): held as found, with nothing to compare until it runs.
away, err := managerAway(ctx, sys, res.Scope, res.User, run)
if err != nil {
return out, h, err
}
if away != "" {
detail = "its unit was found in " + res.User + "'s own manager, which is not running; " +
"kept as found until " + module + " is taken"
break
}
run = managerFor(res.Scope, res.User, run)
}
state, err := sys.ServiceState(ctx, run, res.Unit) state, err := sys.ServiceState(ctx, run, res.Unit)
switch { switch {
case err != nil && !already: case err != nil && !already:
+126
View File
@@ -0,0 +1,126 @@
package apply
import (
"context"
"os"
osuser "os/user"
"path/filepath"
"sort"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// Defends novox/hq ADR 0182 and to-be 41: a parent the host makes inside an owner's home is the
// owner's, one that was there is held as found, and one outside the home is made as before.
// aHome gives the account running the test a home in a directory the test owns, and writes down
// every directory the host gives to whom. The account's own name, so what the host chowns resolves
// without being root; the record, so what was given is told apart from what was merely made.
func aHome(t *testing.T) (home, owner string, given map[string]string) {
t.Helper()
me, err := osuser.Current()
if err != nil {
t.Skip("no current user to own anything")
}
home = t.TempDir()
given = map[string]string{}
wasHome, wasOwn := homeOf, ownMade
homeOf = func(name string) (string, error) {
if name == me.Username {
return home, nil
}
return wasHome(name)
}
ownMade = func(path, owner string) error {
given[path] = owner
return wasOwn(path, owner)
}
t.Cleanup(func() { homeOf, ownMade = wasHome, wasOwn })
return home, me.Username, given
}
func givenPaths(given map[string]string) []string {
var paths []string
for p := range given {
paths = append(paths, p)
}
sort.Strings(paths)
return paths
}
func TestAFileUnderAHomeGivesTheParentsItMadeToItsOwner(t *testing.T) {
home, owner, given := aHome(t)
target := filepath.Join(home, ".config", "mesh", "environment.sh")
d := parse(t, `{"declaration":1,"resources":[{"id":"shell.env","type":"file","path":"`+target+
`","content":"export A=1\n","owner":"`+owner+`"}]}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared,
noServices, nil, nil); err != nil {
t.Fatal(err)
}
want := []string{filepath.Join(home, ".config"), filepath.Join(home, ".config", "mesh")}
if got := givenPaths(given); strings.Join(got, ",") != strings.Join(want, ",") {
t.Errorf("given to the owner: %v, want %v", got, want)
}
for _, p := range want {
if given[p] != owner {
t.Errorf("%s given to %q", p, given[p])
}
}
}
func TestAnArchiveUnderAHomeGivesTheParentsItMadeToItsOwner(t *testing.T) {
home, owner, given := aHome(t)
body, digest := anArchive(t, map[string]string{"p10k.zsh": "theme"})
target := filepath.Join(home, ".local", "share", "powerlevel10k")
d := declare(t, `{"id":"shell.theme","type":"archive","source":"`+serving(t, body)+
`","digest":"`+digest+`","path":"`+target+`","owner":"`+owner+`"}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared,
noServices, nil, nil); err != nil {
t.Fatal(err)
}
for _, p := range []string{filepath.Join(home, ".local"), filepath.Join(home, ".local", "share")} {
if given[p] != owner {
t.Errorf("%s, made by the host, was not given to the owner: %v", p, givenPaths(given))
}
}
}
func TestAParentThatWasThereIsHeldAsFound(t *testing.T) {
home, owner, given := aHome(t)
config := filepath.Join(home, ".config")
if err := os.Mkdir(config, 0o700); err != nil {
t.Fatal(err)
}
target := filepath.Join(config, "mesh", "environment.sh")
d := parse(t, `{"declaration":1,"resources":[{"id":"shell.env","type":"file","path":"`+target+
`","content":"export A=1\n","owner":"`+owner+`"}]}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared,
noServices, nil, nil); err != nil {
t.Fatal(err)
}
if _, touched := given[config]; touched {
t.Error("a parent that was already there was given to the owner")
}
if info, _ := os.Stat(config); info.Mode().Perm() != 0o700 {
t.Errorf("a parent that was already there changed mode: %o", info.Mode().Perm())
}
if given[filepath.Join(config, "mesh")] != owner {
t.Errorf("the parent the host made was not given to the owner: %v", givenPaths(given))
}
}
func TestAParentOutsideTheHomeIsMadeAsBefore(t *testing.T) {
_, owner, given := aHome(t)
target := filepath.Join(t.TempDir(), "var", "lib", "module", "settings.conf")
d := parse(t, `{"declaration":1,"resources":[{"id":"module.conf","type":"file","path":"`+target+
`","content":"a=1\n","owner":"`+owner+`"}]}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared,
noServices, nil, nil); err != nil {
t.Fatal(err)
}
if len(given) != 0 {
t.Errorf("parents outside the owner's home were given to it: %v", givenPaths(given))
}
}
+1 -1
View File
@@ -132,7 +132,7 @@ func applyInto(r *declaration.File, previous store.Applied) (Outcome, error) {
mode = m mode = m
} }
} }
if err := os.MkdirAll(filepath.Dir(r.Path), 0o755); err != nil { if err := makeDirs(filepath.Dir(r.Path), 0o755, r.Owner); err != nil {
return out, err return out, err
} }
if err := writeAtomically(r.Path, want, mode); err != nil { if err := writeAtomically(r.Path, want, mode); err != nil {
+78
View File
@@ -134,3 +134,81 @@ func TestALeftOutModuleIsNeitherRemovedNorForgotten(t *testing.T) {
t.Fatalf("keeping the left-out module's container was not said: %+v", report.Outcomes) t.Fatalf("keeping the left-out module's container was not said: %+v", report.Outcomes)
} }
} }
// A container's capabilities reach the runtime and are part of its spec (novox/hq ADR 0170).
func TestACapabilityReachesTheRuntimeAndTheSpec(t *testing.T) {
var ran []string
run := func(_ context.Context, name string, args ...string) (string, error) {
if name != "docker" {
return "", errors.New("not installed")
}
switch args[0] {
case "info":
return "29.0.0\n", nil
case "container":
return "false\t\n", errors.New("no such container")
case "run":
ran = args
return "deadbeef\n", nil
}
return "", nil
}
d := parseTrusted(t, `{"declaration":1,"resources":[
{"id":"fw","type":"container","name":"fw","image":"`+pinned+`","network":"host","capabilities":["NET_ADMIN"]}
]}`)
_, _, _ = Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil)
granted := false
for i, a := range ran {
if a == "--cap-add" && i+1 < len(ran) && ran[i+1] == "NET_ADMIN" {
granted = true
}
}
if !granted {
t.Fatalf("the capability was not granted: %v", ran)
}
with := d.Resources[0].(*declaration.Container)
without := *with
without.Capabilities = nil
if containerSpec(with, inputs{}) == containerSpec(&without, inputs{}) {
t.Fatal("a capability is not part of the container's spec")
}
}
// A package may be declared absent (novox/hq ADR 0180): removed when it is installed, read back,
// left alone when it is not.
func TestAPackageDeclaredAbsentIsRemovedWhenPresentAndLeftWhenNot(t *testing.T) {
installed := true
var ran []string
run := func(_ context.Context, name string, args ...string) (string, error) {
ran = append(ran, name+" "+strings.Join(args, " "))
if name != "pacman" {
return "", nil
}
switch args[0] {
case "-Q":
if args[1] == "pacman" || installed {
return args[1] + " 1.0\n", nil
}
return "", errors.New("package not found")
case "-R":
installed = false
}
return "", nil
}
d := parseTrusted(t, `{"declaration":1,"resources":[{"id":"front-end","type":"package","package":"ufw","absent":true}]}`)
report, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil)
if err != nil {
t.Fatal(err)
}
if report.Outcomes[0].Action != "removed" || !strings.Contains(strings.Join(ran, "\n"), "pacman -R --noconfirm ufw") {
t.Fatalf("an installed package declared absent was not removed: %+v\n%v", report.Outcomes[0], ran)
}
ran = nil
report, _, err = Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil)
if err != nil {
t.Fatal(err)
}
if report.Outcomes[0].Action != "unchanged" || strings.Contains(strings.Join(ran, "\n"), "-R") {
t.Fatalf("a package already absent was touched: %+v\n%v", report.Outcomes[0], ran)
}
}
+43
View File
@@ -0,0 +1,43 @@
package apply
import (
"context"
"strings"
"testing"
"github.com/novox/mesh-host/internal/declaration"
"github.com/novox/mesh-host/internal/store"
)
// A container declared to log to the journal is run with the journal as its log driver, and the
// place it logs is part of its spec, so moving it recreates the container (novox/hq ADR 0179).
func TestAContainerLoggingToTheJournalIsRunThatWayAndRecreatedWhenMoved(t *testing.T) {
pinned := "postgres@sha256:" + strings.Repeat("a", 64)
var ran []string
run := func(_ context.Context, cmd string, args ...string) (string, error) {
if cmd == "docker" && len(args) > 0 && args[0] == "run" {
ran = args
return "deadbeef\n", nil
}
return "", nil
}
d := parseTrusted(t, `{"declaration":1,"resources":[
{"id":"front","type":"container","name":"front","image":"`+pinned+`","logging":"journald"}
]}`)
_, _, _ = Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil)
sent := false
for i, a := range ran {
if a == "--log-driver" && i+1 < len(ran) && ran[i+1] == "journald" {
sent = true
}
}
if !sent {
t.Fatalf("the container's output was not sent to the journal: %v", ran)
}
with := d.Resources[0].(*declaration.Container)
without := *with
without.Logging = ""
if containerSpec(with, inputs{}) == containerSpec(&without, inputs{}) {
t.Fatal("where a container logs is not part of its spec, so moving it would not recreate it")
}
}
+238
View File
@@ -0,0 +1,238 @@
package apply
import (
"context"
"errors"
"strings"
"sync"
"testing"
"time"
"github.com/novox/mesh-host/internal/declaration"
)
// A scheduled step may hold its module's own containers still while it runs (novox/hq ADR 0189,
// issue 108).
//
// What it exists for: the artifact store's collector walks the storage and requires every writer
// stopped. A run-once step runs beside containers and a scheduled one is the same container again,
// so the mesh had no way to say it — which is why the store it inherited has never collected
// anything. The risk the field brings is one shape only: a window that opens and never closes.
// Every test here is about that shape.
// windowRun records the order of stop / run / start, which is the whole of what is being asserted.
type windowRun struct {
mu sync.Mutex
order []string
failAt string // the arg[0] that should fail ("run" makes the step fail)
wontGo string // a container name that refuses to start again
}
func (w *windowRun) run(_ context.Context, _ string, args ...string) (string, error) {
w.mu.Lock()
defer w.mu.Unlock()
switch args[0] {
case "info":
return "27.0\n", nil
case "stop", "start":
w.order = append(w.order, args[0]+" "+args[1])
if args[0] == "start" && args[1] == w.wontGo {
return "", errors.New("the runtime refused")
}
case "run":
w.order = append(w.order, "run")
if w.failAt == "run" {
return "", errors.New("the step exited non-zero")
}
}
return "", nil
}
func (w *windowRun) seen() []string {
w.mu.Lock()
defer w.mu.Unlock()
return append([]string{}, w.order...)
}
// aStoreWithACollector is a module in the shape distribution has: a server that must not be
// writing, and a nightly step that walks its storage with the server held still.
func aStoreWithACollector(t *testing.T) *declaration.Declaration {
t.Helper()
return parseTrusted(t, `{"declaration":1,"resources":[
{"id":"store","type":"container","name":"mesh-registry","image":"`+pinned+`"},
{"id":"collect","type":"container","name":"mesh-registry-collect","image":"`+pinned+`",
"schedule":"30 3 * * *","while-stopped":["store"]}
]}`)
}
func fireOnce(t *testing.T, d *declaration.Declaration, w *windowRun) {
t.Helper()
clock := &fixedClock{now: time.Date(2026, 10, 2, 3, 29, 0, 0, time.UTC)}
s := NewScheduler(clock, w.run, func(string) {})
s.Sync(d, nil)
s.Advance(context.Background(), time.Date(2026, 10, 2, 3, 30, 5, 0, time.UTC))
s.Wait()
}
func TestAScheduledStepHoldsItsModulesContainerStillAndStartsItAgain(t *testing.T) {
w := &windowRun{}
fireOnce(t, aStoreWithACollector(t), w)
got := w.seen()
want := []string{"stop mesh-registry", "run", "start mesh-registry"}
var kept []string
for _, line := range got {
if strings.HasPrefix(line, "stop mesh-registry-collect") {
// Clearing the step's own exited container by name; not part of the window.
continue
}
kept = append(kept, line)
}
if len(kept) != len(want) {
t.Fatalf("the window was not stop, run, start: %v", got)
}
for i := range want {
if kept[i] != want[i] {
t.Fatalf("the window was %v, want %v", kept, want)
}
}
}
// The one that matters: a step that fails must leave the service running.
func TestAFailedStepStillClosesTheWindow(t *testing.T) {
w := &windowRun{failAt: "run"}
fireOnce(t, aStoreWithACollector(t), w)
var started bool
for _, line := range w.seen() {
if line == "start mesh-registry" {
started = true
}
}
if !started {
t.Fatalf("the step failed and the container it held still was never started again: %v", w.seen())
}
}
// A container that will not come back is said loudly: it is down, and nothing else notices until
// the next apply compares it.
func TestAContainerThatWillNotStartAgainIsSaidLoudly(t *testing.T) {
w := &windowRun{wontGo: "mesh-registry"}
var said []string
clock := &fixedClock{now: time.Date(2026, 10, 2, 3, 29, 0, 0, time.UTC)}
s := NewScheduler(clock, w.run, func(line string) { said = append(said, line) })
s.Sync(aStoreWithACollector(t), nil)
s.Advance(context.Background(), time.Date(2026, 10, 2, 3, 30, 5, 0, time.UTC))
s.Wait()
var loud bool
for _, line := range said {
if strings.Contains(line, "WILL NOT START AGAIN") && strings.Contains(line, "mesh-registry") {
loud = true
}
}
if !loud {
t.Fatalf("a service left stopped by a maintenance window was not said loudly: %v", said)
}
}
// Several containers come back in the reverse of the order they were stopped: a module names the
// dependant first, and starting it before what it depends on is not bringing it back.
func TestTheWindowClosesInTheReverseOfTheOrderItOpened(t *testing.T) {
d := parseTrusted(t, `{"declaration":1,"resources":[
{"id":"web","type":"container","name":"web","image":"`+pinned+`"},
{"id":"db","type":"container","name":"db","image":"`+pinned+`"},
{"id":"collect","type":"container","name":"collect","image":"`+pinned+`",
"schedule":"30 3 * * *","while-stopped":["web","db"]}
]}`)
w := &windowRun{}
fireOnce(t, d, w)
var stops, starts []string
for _, line := range w.seen() {
switch {
case line == "stop web" || line == "stop db":
stops = append(stops, line)
case strings.HasPrefix(line, "start "):
starts = append(starts, line)
}
}
if len(stops) != 2 || stops[0] != "stop web" || stops[1] != "stop db" {
t.Fatalf("stopped in %v, want the order the step named them", stops)
}
if len(starts) != 2 || starts[0] != "start db" || starts[1] != "start web" {
t.Fatalf("started in %v, want the reverse", starts)
}
}
// And the refusals, each for what it says rather than that it says something.
func TestAMaintenanceWindowIsRefusedWhereItCannotMean(t *testing.T) {
for _, c := range []struct{ name, body, says string }{
{
"a window with no schedule",
`{"id":"collect","type":"container","name":"c","image":"` + pinned + `","while-stopped":["store"]}`,
"needs a schedule",
},
{
"a window naming itself",
`{"id":"collect","type":"container","name":"c","image":"` + pinned + `","schedule":"30 3 * * *","while-stopped":["collect"]}`,
"this step itself",
},
{
"a window naming something that is not a container here",
`{"id":"collect","type":"container","name":"c","image":"` + pinned + `","schedule":"30 3 * * *","while-stopped":["elsewhere"]}`,
"no container by that id",
},
} {
_, err := declaration.ParseTrusted([]byte(`{"declaration":1,"resources":[` + c.body + `]}`))
if err == nil {
t.Errorf("%s was accepted", c.name)
continue
}
if !strings.Contains(err.Error(), c.says) {
t.Errorf("%s: the refusal does not say %q: %v", c.name, c.says, err)
}
}
}
// A changed window is a changed declaration, and the install says so.
//
// The cadence already works this way: "a changed schedule is a changed spec — the marker moves and
// the install is reported updated and re-established" (containerSpec). Which containers are held
// still for the run is the same kind of statement, and a declaration that changed it while the
// machine reported no change would be a machine quietly running the old window.
func TestAChangedWindowMovesTheSpec(t *testing.T) {
one := parseTrusted(t, `{"declaration":1,"resources":[
{"id":"store","type":"container","name":"mesh-registry","image":"`+pinned+`"},
{"id":"other","type":"container","name":"other","image":"`+pinned+`"},
{"id":"collect","type":"container","name":"collect","image":"`+pinned+`",
"schedule":"30 3 * * *","while-stopped":["store"]}
]}`)
two := parseTrusted(t, `{"declaration":1,"resources":[
{"id":"store","type":"container","name":"mesh-registry","image":"`+pinned+`"},
{"id":"other","type":"container","name":"other","image":"`+pinned+`"},
{"id":"collect","type":"container","name":"collect","image":"`+pinned+`",
"schedule":"30 3 * * *","while-stopped":["store","other"]}
]}`)
stepOf := func(d *declaration.Declaration) *declaration.Container {
for _, r := range d.Resources {
if c, ok := r.(*declaration.Container); ok && c.ID == "collect" {
return c
}
}
t.Fatal("no step in the fixture")
return nil
}
if containerSpec(stepOf(one), inputs{}) == containerSpec(stepOf(two), inputs{}) {
t.Fatal("the window changed and the spec did not; the machine would report no change " +
"and keep holding the containers it held yesterday")
}
// And a container with no window is untouched by the field existing at all.
plain := parseTrusted(t, `{"declaration":1,"resources":[
{"id":"store","type":"container","name":"mesh-registry","image":"`+pinned+`"}
]}`)
spec := containerSpec(plain.Resources[0].(*declaration.Container), inputs{})
if strings.Contains(spec, "while-stopped") || strings.Contains(spec, "held") {
t.Errorf("an ordinary container's spec mentions a field it does not set:\n%s", spec)
}
}
+51 -12
View File
@@ -54,20 +54,53 @@ func foundFirewall(ctx context.Context, d *declaration.Declaration, known *store
return kind, nil return kind, nil
} }
// retireFirewall disables the found firewall once a converged declaration has applied cleanly, // retireFirewall keeps the found firewall retired on a converged machine (novox/hq ADR 0100, ADR
// which is when the mesh's derived filter has taken its place. Disabled, never flushed: its // 0168): disabled, never flushed, its configuration left on disk for a return to adopted, and the
// configuration stays on disk for a return to adopted, and the container runtime's rules are not // container runtime's rules not its to take.
// its to take. //
// **Convergence is a state the host keeps, not a step it takes once.** Every converged apply reads
// whether the front end is in force; enabled again by a package, a boot or a hand, it is retired
// again and said. The record says how it came to be inactive — the mesh disabled it, or a reconcile
// found it so — and the two are never confused: a flip that did not take, followed by a hand that
// did, used to be recorded as the mesh's doing (issue 143).
// //
// Only a declaration from the mesh converges a node. A carried bundle never says a node is adopted // Only a declaration from the mesh converges a node. A carried bundle never says a node is adopted
// — it cannot — so its silence is not the controller's word that the node was converged, and an // — it cannot — so its silence is not the controller's word that the node was converged, and an
// adopted node re-applying its bundle keeps the firewall it was found with. // adopted node re-applying its bundle keeps the firewall it was found with.
//
// Returned is what this apply did about the found firewall, for the report; empty when the machine
// has none or is not converged.
func retireFirewall(ctx context.Context, d *declaration.Declaration, origin string, known *store.State, func retireFirewall(ctx context.Context, d *declaration.Declaration, origin string, known *store.State,
run Runner, log func(string)) error { run Runner, log func(string)) (string, error) {
rec := known.Firewall rec := known.Firewall
if origin != store.OriginDeclared || d.Adoption != nil || rec == nil || rec.Kind != string(firewall.UFW) || !rec.WasActive || if origin != store.OriginDeclared || d.Adoption != nil || rec == nil || rec.Kind != string(firewall.UFW) || !rec.WasActive {
rec.DisabledByMesh { return "", nil
return nil }
if !firewall.Installed(ctx, run) {
// Uninstalled (novox/hq ADR 0180): retired for good, by the module that replaced it. Said
// once, and nothing is asked of a command that is not there.
if rec.RetiredBy != firewall.RetiredRemoved {
rec.RetiredBy = firewall.RetiredRemoved
log(" the found firewall (ufw) is no longer installed; the mesh's filter is what filters this machine")
return "removed: ufw is no longer installed; the mesh's filter is what filters this machine", nil
}
return "", nil
}
active := firewall.Active(ctx, run)
if !active && !(rec.Forward != nil && !rec.DisabledByMesh) {
// Inactive, and either the mesh's doing already or nobody's recorded here: said as found,
// never as done (issue 143's second fault). A retirement the mesh began and did not finish —
// the forward policy recorded, ufw down, the restore failed — is the one inactive state that
// is still the mesh's to complete, below.
if rec.RetiredBy == "" {
if rec.DisabledByMesh {
rec.RetiredBy = firewall.RetiredByMesh
} else {
rec.RetiredBy = firewall.RetiredFoundSo
log(" ufw is inactive on this converged node, and not by the mesh; recorded as found so")
}
}
return "", nil
} }
// **Nothing is retired until what replaces it is in force** (novox/hq ADR 0100). The flip // **Nothing is retired until what replaces it is in force** (novox/hq ADR 0100). The flip
// loads the mesh's derived filter in ufw's place; disabling ufw before that table is actually // loads the mesh's derived filter in ufw's place; disabling ufw before that table is actually
@@ -75,10 +108,10 @@ func retireFirewall(ctx context.Context, d *declaration.Declaration, origin stri
// no filter at all. // no filter at all.
loaded, err := firewall.MeshTableLoaded(ctx, run) loaded, err := firewall.MeshTableLoaded(ctx, run)
if err != nil { if err != nil {
return err return "", err
} }
if !loaded { if !loaded {
return fmt.Errorf("this node is converged and the mesh's own filter (table %s) is not loaded on "+ return "", fmt.Errorf("this node is converged and the mesh's own filter (table %s) is not loaded on "+
"this machine, so ufw was left in force: retiring it would leave the machine filtering "+ "this machine, so ufw was left in force: retiring it would leave the machine filtering "+
"nothing. Assign a filter module to this node, or return it to adopted", firewall.MeshTable) "nothing. Assign a filter module to this node, or return it to adopted", firewall.MeshTable)
} }
@@ -88,11 +121,17 @@ func retireFirewall(ctx context.Context, d *declaration.Declaration, origin stri
rec.Forward = firewall.ForwardPolicies(ctx, run) rec.Forward = firewall.ForwardPolicies(ctx, run)
} }
if err := firewall.Disable(ctx, run, rec.Forward); err != nil { if err := firewall.Disable(ctx, run, rec.Forward); err != nil {
return err return "", err
} }
again := rec.DisabledByMesh || rec.RetiredBy != ""
rec.DisabledByMesh = true rec.DisabledByMesh = true
rec.RetiredBy = firewall.RetiredByMesh
if again {
log(" disabled ufw again: it had been enabled since the mesh retired it; this node is converged and filtered by the mesh")
return "disabled again: ufw had been enabled since the mesh retired it", nil
}
log(" disabled ufw: this node is converged and filtered by the mesh; ufw's configuration is left on disk") log(" disabled ufw: this node is converged and filtered by the mesh; ufw's configuration is left on disk")
return nil return "disabled: this node is converged and filtered by the mesh; ufw's configuration is left on disk", nil
} }
// applyOpening makes one opening true through the firewall found here. // applyOpening makes one opening true through the firewall found here.
+53 -2
View File
@@ -8,6 +8,7 @@ import (
"path/filepath" "path/filepath"
"strings" "strings"
"testing" "testing"
"time"
"github.com/novox/mesh-host/internal/declaration" "github.com/novox/mesh-host/internal/declaration"
"github.com/novox/mesh-host/internal/store" "github.com/novox/mesh-host/internal/store"
@@ -189,14 +190,34 @@ func TestConvergingRetiresTheFoundFirewallAndReturningRestoresIt(t *testing.T) {
} }
} }
// Converged again: nothing more to retire. // Converged again: nothing more to retire — the node asks ufw whether it is in force, which is
// what keeps convergence a state rather than a step taken once (novox/hq ADR 0168), and touches
// nothing else.
u.asked = nil u.asked = nil
if _, state, err = applyWith(t, converged, state, u.run); err != nil { if _, state, err = applyWith(t, converged, state, u.run); err != nil {
t.Fatal(err) t.Fatal(err)
} }
if u.index("ufw") >= 0 { for _, a := range u.asked {
if strings.HasPrefix(a, "ufw") && a != "ufw status" {
t.Errorf("a converged node kept talking to a retired ufw: %v", u.asked) t.Errorf("a converged node kept talking to a retired ufw: %v", u.asked)
} }
}
if state.Firewall.RetiredBy != "mesh" {
t.Errorf("the record does not say the mesh retired it: %+v", state.Firewall)
}
// Enabled again by a hand: retired again, and said.
u.active = true
u.asked = nil
report, state, err := applyWith(t, converged, state, u.run)
if err != nil {
t.Fatal(err)
}
if u.active || u.index("ufw disable") < 0 {
t.Fatalf("ufw enabled again on a converged node was not retired again: %v", u.asked)
}
if !strings.Contains(report.Firewall, "disabled again") {
t.Errorf("retiring it again was not said: %q", report.Firewall)
}
// Returned to adopted: ufw is enabled before the opening is converged through it. // Returned to adopted: ufw is enabled before the opening is converged through it.
u.asked = nil u.asked = nil
@@ -502,3 +523,33 @@ func TestUfwIsNotRetiredUntilTheMeshsOwnFilterIsLoaded(t *testing.T) {
t.Errorf("ufw was not retired once the mesh's filter was loaded: active %v, %+v", u.active, state.Firewall) t.Errorf("ufw was not retired once the mesh's filter was loaded: active %v, %+v", u.active, state.Firewall)
} }
} }
// A front end that is no longer installed is recorded as removed, said once, and asked nothing of
// (novox/hq ADR 0180).
func TestAnUninstalledFrontEndIsRetiredForGood(t *testing.T) {
dir := t.TempDir()
u := &ufwMachine{installed: false, ruleset: "table inet mesh\n"}
known := store.State{Firewall: &store.FoundFirewall{Kind: "ufw", WasActive: true, DisabledByMesh: true,
RetiredBy: "mesh", FoundAt: time.Now()}}
converged := parse(t, `{"declaration":1,"resources":[`+withConf(dir)+`]}`)
report, state, err := applyWith(t, converged, known, u.run)
if err != nil {
t.Fatal(err)
}
if state.Firewall.RetiredBy != "removed" || !strings.Contains(report.Firewall, "no longer installed") {
t.Fatalf("record %+v, said %q", state.Firewall, report.Firewall)
}
u.asked = nil
report, _, err = applyWith(t, converged, state, u.run)
if err != nil {
t.Fatal(err)
}
if report.Firewall != "" {
t.Errorf("said again: %q", report.Firewall)
}
for _, a := range u.asked {
if strings.HasPrefix(a, "ufw") && a != "ufw status" {
t.Errorf("asked something of a front end that is not there: %v", u.asked)
}
}
}
+12 -1
View File
@@ -90,6 +90,17 @@ func Plan(d *declaration.Declaration, known store.State, origin string) []Step {
switch { switch {
case orphan.Stateless: case orphan.Stateless:
step.Verb, step.Why = "forget", "no longer declared; its unit's state was never the mesh's and is left as it is" step.Verb, step.Why = "forget", "no longer declared; its unit's state was never the mesh's and is left as it is"
case orphan.Type == string(declaration.TypeArchive) && store.IsFormer(orphan.ID):
// In removeArchive's words (novox/hq issues 162, 194).
step.Verb, step.Why = "forget", "a former target; left in place, since the version before is what a rollback starts"
case orphan.Type == string(declaration.TypeArchive) && orphan.Unpacked == nil:
step.Verb, step.Why = "forget", "no longer declared; recorded before the host kept what it unpacked, so it is left in place"
case orphan.Type == string(declaration.TypeArchive):
step.Why = "no longer declared; the files it unpacked go, and the directories the host made for them once empty"
case orphan.Type == string(declaration.TypeFile) && orphan.Into == nil && orphan.Kept != "":
// In removeWhole's words (novox/hq ADR 0118): written over, so given back, not deleted.
step.Verb, step.Why = "restore", "no longer declared; the original the mesh wrote over goes back "+
"from "+orphan.Kept+", unless the file was changed since the mesh last wrote it"
case orphan.Type == string(declaration.TypeService): case orphan.Type == string(declaration.TypeService):
// What removal will do, said before it does it (novox/hq ADR 0118), in removeService's // What removal will do, said before it does it (novox/hq ADR 0118), in removeService's
// words. "restore" only where it may stop or disable something — the record cannot say // words. "restore" only where it may stop or disable something — the record cannot say
@@ -98,7 +109,7 @@ func Plan(d *declaration.Declaration, known store.State, origin string) []Step {
// stopped whatever was found. // stopped whatever was found.
f := orphan.Found f := orphan.Found
switch { switch {
case made[orphan.Target]: case made[unitKey(orphan.Scope, orphan.User, orphan.Target)]:
step.Verb, step.Why = "remove", "no longer declared; the mesh wrote its unit file, so it is "+ step.Verb, step.Why = "remove", "no longer declared; the mesh wrote its unit file, so it is "+
"stopped and disabled at boot before that file goes" "stopped and disabled at boot before that file goes"
case f == nil: case f == nil:
+47 -7
View File
@@ -77,11 +77,26 @@ func applyProcess(ctx context.Context, r *declaration.Process, run Runner,
// Everything about it is as declared. Still asked whether it is RUNNING, because a // Everything about it is as declared. Still asked whether it is RUNNING, because a
// declaration that is satisfied by a record rather than by the machine is how a stopped // declaration that is satisfied by a record rather than by the machine is how a stopped
// service reports success. // service reports success.
if active, err := run(ctx, "systemctl", "is-active", "--quiet", r.Name+".service"); err == nil { // **The record is carried forward, not re-derived.** An unchanged outcome is recorded
// like any other, so one that said nothing about what was written erased the digest; the
// next apply then found no record, re-created the daemon, and the one after that found a
// record again — the node's runtime restarted every other cycle (novox/hq 04-ISSUES/210).
out.wrote = want
// **A step that ran is done, and a schedule is its timer** (novox/hq issue 213). Neither's
// service is meant to be up between runs, so asking whether it is — and starting it when it
// was not — ran a completed step again on every apply, and a scheduled run off its cadence.
if r.RunOnce {
return out, nil
}
unit := r.Name + ".service"
if r.Schedule != "" {
unit = r.Name + ".timer"
}
if active, err := run(ctx, "systemctl", "is-active", "--quiet", unit); err == nil {
_ = active _ = active
return out, nil return out, nil
} }
if _, err := run(ctx, "systemctl", "start", r.Name+".service"); err != nil { if _, err := run(ctx, "systemctl", "start", unit); err != nil {
return out, fmt.Errorf("%s is installed and would not start: %w", r.Name, err) return out, fmt.Errorf("%s is installed and would not start: %w", r.Name, err)
} }
out.Action = "updated" out.Action = "updated"
@@ -110,9 +125,22 @@ func applyProcess(ctx context.Context, r *declaration.Process, run Runner,
// so the machine is not asked to start something that needed a migration that did not happen. // so the machine is not asked to start something that needed a migration that did not happen.
// Nothing is left behind to ask afterwards: the record that it ran is the digest, which is why // Nothing is left behind to ask afterwards: the record that it ran is the digest, which is why
// the identity above includes the command. // the identity above includes the command.
//
// **Run as the unit a daemon would be, once** (novox/hq design 38 WP4c). Run directly, the
// step started in the host's own working directory, without its environment, its environment
// files or its user — `node bootstrap/index.js` resolved from wherever the host ran and was told
// none of the words it was declared with. A oneshot unit carries all four exactly as a daemon's
// does, and starting one waits for it to finish and fails when it fails.
if r.RunOnce { if r.RunOnce {
if _, err := run(ctx, r.Run[0], r.Run[1:]...); err != nil { unit := filepath.Join(unitDir, r.Name+".service")
return out, fmt.Errorf("the %s step did not complete: %w", r.Name, err) if err := os.WriteFile(unit, []byte(unitFor(r)), 0o644); err != nil {
return out, err
}
if _, err := run(ctx, "systemctl", "daemon-reload"); err != nil {
return out, err
}
if _, err := run(ctx, "systemctl", "start", r.Name+".service"); err != nil {
return out, fmt.Errorf("the %s step did not complete (journalctl -u %s.service says why): %w", r.Name, r.Name, err)
} }
out.Action = "created" out.Action = "created"
if previous.Wrote != "" { if previous.Wrote != "" {
@@ -209,9 +237,9 @@ func unitFor(r *declaration.Process) string {
if r.User != "" { if r.User != "" {
fmt.Fprintf(&b, "User=%s\n", r.User) fmt.Fprintf(&b, "User=%s\n", r.User)
} }
fmt.Fprintf(&b, "ExecStart=%s\n", strings.Join(r.Run, " ")) fmt.Fprintf(&b, "ExecStart=%s\n", strings.Join(runFrom(r), " "))
if r.Schedule != "" { if r.Schedule != "" || r.RunOnce {
// Started by its timer and expected to finish. Restarting it would have it run // Started by its timer, or once by the host, and expected to finish. Restarting it would have it run
// continuously between fires, which is the opposite of a schedule. // continuously between fires, which is the opposite of a schedule.
b.WriteString("Type=oneshot\n") b.WriteString("Type=oneshot\n")
b.WriteString("\n") b.WriteString("\n")
@@ -341,3 +369,15 @@ func removeProcess(ctx context.Context, a store.Applied, run Runner) (string, st
} }
return "removed", "stopped; its unit and its bundle removed — the mesh's own code", nil return "removed", "stopped; its unit and its bundle removed — the mesh's own code", nil
} }
// runFrom is the command as the unit runs it. **A command written `./name` is that file in the
// process's own unpacked bundle** (novox/hq ADR 0193): a bundle compiled to a binary runs itself,
// and only the host knows where it unpacked it, while the service manager takes an absolute path or
// a name it finds on its own search path — never one relative to the working directory.
func runFrom(r *declaration.Process) []string {
run := append([]string(nil), r.Run...)
if len(run) > 0 && strings.HasPrefix(run[0], "./") {
run[0] = filepath.Join(daemonRoot, r.Name, strings.TrimPrefix(run[0], "./"))
}
return run
}
+81
View File
@@ -0,0 +1,81 @@
package apply
import (
"context"
"errors"
"os"
"os/user"
"path/filepath"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// A run-once process is a step run where, how and as whom it was declared (novox/hq design 38
// WP4c): its bundle's directory, its environment and environment files, its user. Run directly,
// it started in the host's own directory with none of them.
func TestARunOnceProcessRunsAsItsOneshotUnit(t *testing.T) {
units, bundles := t.TempDir(), t.TempDir()
wasUnits, wasBundles := unitDir, daemonRoot
unitDir, daemonRoot = units, bundles
t.Cleanup(func() { unitDir, daemonRoot = wasUnits, wasBundles })
me, err := user.Current()
if err != nil {
t.Fatal(err)
}
body, digest := anArchive(t, map[string]string{"bootstrap/index.js": "console.log(1)\n"})
var commands []string
fail := false
run := func(ctx context.Context, name string, args ...string) (string, error) {
commands = append(commands, name+" "+strings.Join(args, " "))
if fail && name == "systemctl" && len(args) > 0 && args[0] == "start" {
return "", errors.New("exit status 1")
}
return "", nil
}
d := declare(t, `{"id":"mosquitto.bootstrap","type":"process","name":"mosquitto-bootstrap","source":"`+serving(t, body)+
`","digest":"`+digest+`","run":["node","bootstrap/index.js"],"run-once":true,"user":"`+me.Username+`",`+
`"env":{"MESH_ADMIN":"mesh-admin"},"env-file":["/var/lib/mesh/mosquitto/bootstrap.env"]}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared, run, nil, nil); err != nil {
t.Fatal(err)
}
unit, err := os.ReadFile(filepath.Join(units, "mosquitto-bootstrap.service"))
if err != nil {
t.Fatalf("no unit was written for the step: %v", err)
}
for _, want := range []string{
"WorkingDirectory=" + filepath.Join(bundles, "mosquitto-bootstrap"),
"EnvironmentFile=/var/lib/mesh/mosquitto/bootstrap.env",
`Environment="MESH_ADMIN=mesh-admin"`,
"User=" + me.Username,
"Type=oneshot",
"ExecStart=node bootstrap/index.js",
} {
if !strings.Contains(string(unit), want) {
t.Errorf("the step's unit lacks %q:\n%s", want, unit)
}
}
for _, never := range []string{"Restart=always", "[Install]", "Type=simple"} {
if strings.Contains(string(unit), never) {
t.Errorf("a step's unit says %q:\n%s", never, unit)
}
}
joined := strings.Join(commands, "; ")
if !strings.Contains(joined, "systemctl start mosquitto-bootstrap.service") {
t.Errorf("the step was not started as its unit: %s", joined)
}
if strings.Contains(joined, "node bootstrap/index.js") || strings.Contains(joined, "enable mosquitto-bootstrap") {
t.Errorf("the step was run directly or enabled: %s", joined)
}
// A step that fails fails the apply, and is not recorded as done.
fail = true
d2 := declare(t, `{"id":"mosquitto.bootstrap","type":"process","name":"mosquitto-bootstrap","source":"`+serving(t, body)+
`","digest":"`+digest+`","run":["node","bootstrap/index.js"],"run-once":true,"env":{"MESH_ADMIN":"changed"}}`)
if _, _, err := Apply(context.Background(), archHost(t), d2, store.State{}, store.OriginDeclared, run, nil, nil); err == nil {
t.Error("a step that failed did not fail the apply")
}
}
+70
View File
@@ -256,3 +256,73 @@ func TestAProcessRecordedUnderAPathlikeNameIsRefusedNotRemoved(t *testing.T) {
} }
} }
} }
// novox/hq 04-ISSUES/210: the node's runtime was re-created — and restarted — on every reconcile,
// because the host did not find what it wrote for a process the cycle before. Applying the same
// process declaration twice must do no work the second time.
func TestAProcessAppliedAgainIsUnchangedAndNotRestarted(t *testing.T) {
units, bundles := t.TempDir(), t.TempDir()
wasUnits, wasBundles := unitDir, daemonRoot
unitDir, daemonRoot = units, bundles
t.Cleanup(func() { unitDir, daemonRoot = wasUnits, wasBundles })
body, digest := anArchive(t, map[string]string{"main.js": "console.log(1)\n"})
var commands []string
run := func(ctx context.Context, name string, args ...string) (string, error) {
commands = append(commands, name+" "+strings.Join(args, " "))
return "", nil
}
d := declare(t, `{"id":"node-tools.runtime","type":"process","name":"node-tools","source":"`+serving(t, body)+
`","digest":"`+digest+`","run":["node","main.js"],"env":{"MESH_TOOL_MODULES":"a=/x/index.js"}}`)
first, state, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginDeclared, run, nil, nil)
if err != nil {
t.Fatal(err)
}
if o := outcomeOf(first, "node-tools.runtime"); o.Action != "created" {
t.Fatalf("first apply: %+v, want created", o)
}
rec, ok := state.Find("node-tools.runtime")
if !ok || rec.Wrote == "" {
t.Fatalf("the host did not record what it wrote for the process: %+v", rec)
}
commands = nil
again, state, err := Apply(context.Background(), archHost(t), d, state, store.OriginDeclared, run, nil, nil)
if err != nil {
t.Fatal(err)
}
if o := outcomeOf(again, "node-tools.runtime"); o.Action != "unchanged" {
t.Errorf("second apply: %+v, want unchanged", o)
}
for _, c := range commands {
if strings.Contains(c, "restart") {
t.Errorf("the second apply restarted the process: %v", commands)
}
}
// And the record survives an unchanged apply: the third cycle is unchanged too. This is the
// cycle the live mesh showed — created, unchanged, created — before the record was carried.
if rec, _ := state.Find("node-tools.runtime"); rec.Wrote == "" {
t.Fatalf("an unchanged apply dropped the digest from the record: %+v", rec)
}
commands = nil
third, _, err := Apply(context.Background(), archHost(t), d, state, store.OriginDeclared, run, nil, nil)
if err != nil {
t.Fatal(err)
}
if o := outcomeOf(third, "node-tools.runtime"); o.Action != "unchanged" {
t.Errorf("third apply: %+v, want unchanged", o)
}
}
// novox/hq ADR 0193: a bundle compiled to a binary runs itself — `./name` is that file in the
// process's own unpacked bundle, made absolute because the service manager takes nothing relative.
func TestAProcessRunsItsOwnBundlesBinary(t *testing.T) {
unit := unitFor(&declaration.Process{Name: "node-tools", Run: []string{"./node-tools", "serve"}})
if !strings.Contains(unit, "ExecStart="+filepath.Join(daemonRoot, "node-tools", "node-tools")+" serve\n") {
t.Errorf("the binary is not run from its bundle:\n%s", unit)
}
other := unitFor(&declaration.Process{Name: "x", Run: []string{"node", "src/main.js"}})
if !strings.Contains(other, "ExecStart=node src/main.js\n") {
t.Errorf("a command found on the path was changed:\n%s", other)
}
}
+86 -1
View File
@@ -65,6 +65,29 @@ type scheduledJob struct {
container *declaration.Container container *declaration.Container
next time.Time // the next minute at which it is due next time.Time // the next minute at which it is due
running bool // a run is in flight — the next due run is skipped rather than stacked running bool // a run is in flight — the next due run is skipped rather than stacked
// hold is the runtime names of the containers held still for the duration of a run, in the
// order the step named them (novox/hq ADR 0189).
hold []string
}
// heldStillFor is the runtime names of the containers a step holds still, resolved from ids.
func heldStillFor(step *declaration.Container, d *declaration.Declaration) []string {
if len(step.WhileStopped) == 0 {
return nil
}
byID := map[string]string{}
for _, r := range d.Resources {
if c, ok := r.(*declaration.Container); ok {
byID[c.Identity()] = c.Name
}
}
out := make([]string, 0, len(step.WhileStopped))
for _, id := range step.WhileStopped {
if name := byID[id]; name != "" {
out = append(out, name)
}
}
return out
} }
// NewScheduler builds a scheduler. A nil clock is the system clock; a nil log says nothing. // NewScheduler builds a scheduler. A nil clock is the system clock; a nil log says nothing.
@@ -123,15 +146,20 @@ func (s *Scheduler) Sync(d *declaration.Declaration, held map[string]bool) {
// on its cadence rather than staying running to be restarted — so its identity cannot // on its cadence rather than staying running to be restarted — so its identity cannot
// depend on another resource's content and there is nothing to pass. // depend on another resource's content and there is nothing to pass.
spec := containerSpec(c, inputs{}) spec := containerSpec(c, inputs{})
// The runtime stops containers by name; the declaration names them by id. Resolved here,
// against the declaration this job was armed from, so a fire never has to look anything up
// (novox/hq ADR 0189). The parser has already refused an id that is not a container here.
hold := heldStillFor(c, d)
if existing := s.jobs[c.Identity()]; existing != nil && existing.spec == spec { if existing := s.jobs[c.Identity()]; existing != nil && existing.spec == spec {
// Unchanged: keep where it is in its cadence, refresh the declaration pointer only. // Unchanged: keep where it is in its cadence, refresh the declaration pointer only.
existing.container = c existing.container = c
existing.hold = hold
continue continue
} }
// New or changed: arm it for the next due minute after now. // New or changed: arm it for the next due minute after now.
next, _ := cron.Next(s.clock.Now()) next, _ := cron.Next(s.clock.Now())
s.jobs[c.Identity()] = &scheduledJob{ s.jobs[c.Identity()] = &scheduledJob{
id: c.Identity(), spec: spec, cron: cron, container: c, next: next, id: c.Identity(), spec: spec, cron: cron, container: c, next: next, hold: hold,
} }
} }
@@ -202,6 +230,15 @@ func (s *Scheduler) fire(ctx context.Context, j *scheduledJob) {
return return
} }
// **The window opens here and closes in the defer, whatever happens** (novox/hq ADR 0189).
// Deferred before the first stop so a panic, a failing step or a step that runs long all end
// the same way: the service running. The one real risk of this field is a window that never
// closes, and the only defence against it is that closing is not conditional on anything.
if len(j.hold) > 0 {
defer s.letRun(ctx, cri, j)
s.holdStill(ctx, cri, j)
}
// A container by this name left exited by the previous run would collide with --name. Removing // A container by this name left exited by the previous run would collide with --name. Removing
// one that is not there is the state we want, so its error is ignored — the same as run-once. // one that is not there is the state we want, so its error is ignored — the same as run-once.
_, _ = s.run(ctx, cri, "rm", "-f", j.container.Name) _, _ = s.run(ctx, cri, "rm", "-f", j.container.Name)
@@ -260,3 +297,51 @@ func (s *Scheduler) Run(ctx context.Context) {
} }
} }
} }
// holdStill stops the containers this step runs instead of, in the order it named them.
//
// A stop that fails is said and not fatal. The step runs anyway: for the case this exists for —
// a collector walking storage nothing must be writing to — a writer that would not stop is worth
// knowing about, and refusing to run would mean the work never happens and the log says nothing
// new each night. What must not be skipped is the restart, and it is not: it is deferred.
func (s *Scheduler) holdStill(ctx context.Context, cri string, j *scheduledJob) {
for _, name := range j.hold {
if _, err := s.run(ctx, cri, "stop", name); err != nil {
s.log(fmt.Sprintf("scheduled step %s: could not stop %s for the run: %v", j.id, name, err))
continue
}
s.log(fmt.Sprintf("scheduled step %s: %s held still for the run", j.id, name))
}
}
// letRun starts them again, in the reverse of the order they were stopped, and says so loudly if
// one does not come back.
//
// **Reverse order**, because stopping walks a dependency the other way: a module that holds two
// containers still names the one that depends on the other first, and bringing them back the same
// way would start a dependant before what it depends on.
//
// Given its own context, because this runs in a defer and the one the run used may already be
// cancelled — a host shutting down mid-window would otherwise leave the service stopped, which is
// precisely the outcome this field must never have.
func (s *Scheduler) letRun(_ context.Context, cri string, j *scheduledJob) {
ctx, cancel := context.WithTimeout(context.Background(), closingWindow)
defer cancel()
for i := len(j.hold) - 1; i >= 0; i-- {
name := j.hold[i]
if _, err := s.run(ctx, cri, "start", name); err != nil {
// Said as loudly as this host says anything: a service the mesh stopped for a
// maintenance window and could not start again is down, and nothing else will notice
// until the next apply compares it.
s.log(fmt.Sprintf(
"scheduled step %s: %s was held still for the run and WILL NOT START AGAIN: %v",
j.id, name, err))
continue
}
s.log(fmt.Sprintf("scheduled step %s: %s running again", j.id, name))
}
}
// closingWindow is how long the host will spend putting back what it stopped. Generous: this is
// the half that must not be given up on.
const closingWindow = 5 * time.Minute
+100
View File
@@ -0,0 +1,100 @@
package apply
import (
"context"
"os"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// A daemon that reads a configuration it refuses dies a fraction of a second after the service
// manager has reported it started — fail2ban took 221 milliseconds on the control node the day this
// was written. One read back catches nothing: the unit is active at that instant. The mesh reported
// the service "restarted" while every ban on two public machines was gone, and every check passed
// (novox/hq ADR 0184). The host looks again, after the moment in which that happens.
func TestAServiceThatDiesJustAfterItsRestartIsNotReportedRestarted(t *testing.T) {
serviceSettle = 0
// Alive at the first look after starting, dead at the second — the shape of a daemon that
// refuses what it was just given.
started, looks := false, 0
run := func(ctx context.Context, name string, args ...string) (string, error) {
if args[0] == "show" {
state := "active"
if started {
if looks++; looks >= 2 {
state = "failed"
}
}
return "LoadState=loaded\nActiveState=" + state + "\n", nil
}
if args[0] == "start" {
started = true
}
return "", nil
}
dir := t.TempDir()
d := parse(t, `{"declaration":1,"resources":[
{"id":"conf","type":"file","path":"`+dir+`/jail.conf","content":"[sshd]\n","mode":"0644"},
{"id":"run","type":"service","unit":"fail2ban.service","state":"running","restart-on":["conf"]}
]}`)
_, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil)
if err == nil {
t.Fatal("a service that died just after being restarted was reported as restarted")
}
if !strings.Contains(err.Error(), "fail2ban.service") || !strings.Contains(err.Error(), "is stopped") {
t.Errorf("the failure does not name the unit and what it is now: %v", err)
}
if looks < 2 {
t.Errorf("the host looked at the unit %d time(s) after starting it; it must look again", looks)
}
}
// A unit still coming up reads as running at both looks and is accepted: the second look is for a
// unit that WAS running and is not any more, never a wait for a slow one to finish starting.
func TestAUnitStillStartingIsNotAFailure(t *testing.T) {
serviceSettle = 0
run := func(ctx context.Context, name string, args ...string) (string, error) {
if args[0] == "show" {
return "LoadState=loaded\nActiveState=activating\n", nil
}
return "", nil
}
d := parse(t, `{"declaration":1,"resources":[
{"id":"s","type":"service","unit":"slow.service","state":"running"}
]}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil); err != nil {
t.Fatalf("a unit still starting was reported as a failure: %v", err)
}
}
// And a service the declaration asks to be stopped is not waited on at all.
func TestAServiceAskedToStopIsNotWaitedOn(t *testing.T) {
serviceSettle = 0
shows := 0
run := func(ctx context.Context, name string, args ...string) (string, error) {
if args[0] == "show" {
shows++
if shows == 1 {
return "LoadState=loaded\nActiveState=active\n", nil
}
return "LoadState=loaded\nActiveState=inactive\n", nil
}
return "", nil
}
d := parse(t, `{"declaration":1,"resources":[
{"id":"s","type":"service","unit":"off.service","state":"stopped"}
]}`)
if _, _, err := Apply(context.Background(), archHost(t), d, store.State{}, store.OriginCarried, run, nil, nil); err != nil {
t.Fatalf("stopping a service was reported as a failure: %v", err)
}
}
// The settle between a unit's two read-backs is a real pause on a machine and nothing in a test:
// no test here drives a service manager that takes time, so paying it would only slow the suite
// (novox/hq ADR 0184).
func TestMain(m *testing.M) {
serviceSettle = 0
os.Exit(m.Run())
}
+13
View File
@@ -0,0 +1,13 @@
package apply
// ForTests points where the host writes units and unpacks daemons at directories a test owns, and
// waits nothing between its looks at a unit, until the returned function puts them back. For tests
// in other packages that apply a process — the bootstrap's, which checks that what genesis raises is
// what the controller's process takes over (novox/hq issue 223). Nothing outside a test calls it.
func ForTests(units, daemons string) (restore func()) {
wasUnits, wasDaemons, wasSettle, wasHandover := unitDir, daemonRoot, serviceSettle, handoverSettle
unitDir, daemonRoot, serviceSettle, handoverSettle = units, daemons, 0, 0
return func() {
unitDir, daemonRoot, serviceSettle, handoverSettle = wasUnits, wasDaemons, wasSettle, wasHandover
}
}
+292 -18
View File
@@ -2,6 +2,7 @@ package apply
import ( import (
"context" "context"
"errors"
"fmt" "fmt"
"os" "os"
osuser "os/user" osuser "os/user"
@@ -10,6 +11,7 @@ import (
"strings" "strings"
"github.com/novox/mesh-host/internal/declaration" "github.com/novox/mesh-host/internal/declaration"
"github.com/novox/mesh-host/internal/store"
"github.com/novox/mesh-host/internal/system" "github.com/novox/mesh-host/internal/system"
) )
@@ -23,14 +25,39 @@ import (
// //
// Reconciling, like everything else here: it is not told whether the user is new. Creating, // Reconciling, like everything else here: it is not told whether the user is new. Creating,
// setting a shell and adding groups are each done only when the machine does not already agree. // setting a shell and adding groups are each done only when the machine does not already agree.
func applyUser(ctx context.Context, sys system.System, r *declaration.User, run Runner) (Outcome, error) { //
// previous is this resource's record, which carries the shell the account had before the mesh
// first changed it, so removal can give it back (novox/hq ADR 0176 §2, issue 228).
func applyUser(ctx context.Context, sys system.System, r *declaration.User, run Runner,
previous store.Applied) (Outcome, error) {
out := begin(r) out := begin(r)
out.Action = "unchanged" out.Action = "unchanged"
// What was found is carried from the record for as long as the resource is recorded — for this
// account only: a declaration that renamed its user says nothing about the new one's shell.
if previous.Shell != nil && previous.Target == r.Name {
kept := *previous.Shell
out.shell = &kept
}
if previous.Linger != nil && previous.Target == r.Name {
kept := *previous.Linger
out.linger = &kept
}
login, exists, err := system.LookUpUser(ctx, system.Runner(run), r.Name) login, exists, err := system.LookUpUser(ctx, system.Runner(run), r.Name)
if err != nil { if err != nil {
return out, err return out, err
} }
// **A shell is refused before anything is touched** (novox/hq issue 228). Refused after the
// account was created or its groups changed, the account would be half the declaration's; a
// refusal fails this resource and leaves the account exactly as it was.
if r.Shell != "" && (!exists || login.Shell != r.Shell) {
if err := system.UsableShell(r.Shell); err != nil {
return out, fmt.Errorf("%q's shell was not set, and the account was left as it is: %w",
r.Name, err)
}
}
if !exists { if !exists {
if err := sys.CreateUser(ctx, system.Runner(run), r.Name, r.Home, r.Shell); err != nil { if err := sys.CreateUser(ctx, system.Runner(run), r.Name, r.Home, r.Shell); err != nil {
return out, err return out, err
@@ -46,26 +73,15 @@ func applyUser(ctx context.Context, sys system.System, r *declaration.User, run
r.Name) r.Name)
} }
out.Action = "created" out.Action = "created"
} if r.Shell != "" {
// No shell from before to give back: the account had none until the mesh made it.
// The shell, only when it differs. Absent means the host asserts nothing — a field that out.shell = &store.LoginShell{Set: login.Shell, Created: true}
// always asserts cannot express "leave it alone", which is the difference between managing a
// machine and taking it over.
if r.Shell != "" && login.Shell != r.Shell {
if err := sys.SetUserShell(ctx, system.Runner(run), r.Name, r.Shell); err != nil {
return out, err
}
if back, _, err := system.LookUpUser(ctx, system.Runner(run), r.Name); err != nil {
return out, err
} else if back.Shell != r.Shell {
return out, fmt.Errorf("set %q's shell to %q and the user database says %q",
r.Name, r.Shell, back.Shell)
}
if out.Action == "unchanged" {
out.Action = "updated"
} }
} }
// Groups before the shell, so that a failure here comes before the shell is changed: a record
// is written only for an apply that worked, and a shell changed by a failed one would be read
// next time as the account's own, and the one it replaced lost.
if len(r.Groups) > 0 { if len(r.Groups) > 0 {
in, err := system.GroupsOf(ctx, system.Runner(run), r.Name) in, err := system.GroupsOf(ctx, system.Runner(run), r.Name)
if err != nil { if err != nil {
@@ -87,9 +103,198 @@ func applyUser(ctx context.Context, sys system.System, r *declaration.User, run
} }
} }
} }
// The shell, only when it differs. Absent means the host asserts nothing — a field that
// always asserts cannot express "leave it alone", which is the difference between managing a
// machine and taking it over.
if r.Shell != "" && login.Shell != r.Shell {
if err := sys.SetUserShell(ctx, system.Runner(run), r.Name, r.Shell); err != nil {
return out, err
}
if back, _, err := system.LookUpUser(ctx, system.Runner(run), r.Name); err != nil {
return out, err
} else if back.Shell != r.Shell {
return out, fmt.Errorf("set %q's shell to %q and the user database says %q",
r.Name, r.Shell, back.Shell)
}
// **What was found is recorded once** (novox/hq ADR 0176 §2). A later change keeps it: what
// is given back is the shell from before the mesh, never the mesh's own earlier choice.
if out.shell == nil {
out.shell = &store.LoginShell{Found: login.Shell}
}
out.shell.Set = r.Shell
if out.Action == "unchanged" {
out.Action = "updated"
}
}
// Lingering last, and only when declared and different: the one change here that starts or
// stops something — the account's manager and every unit in it (novox/hq ADR 0177).
if r.Linger != nil {
changed, err := applyLinger(ctx, sys, r.Name, *r.Linger, run, &out)
if err != nil {
return out, err
}
if changed && out.Action == "unchanged" {
out.Action = "updated"
}
}
return out, nil return out, nil
} }
// lingerer is a service manager that runs a manager per account, and can keep one running with
// nobody logged in (novox/hq ADR 0177).
type lingerer interface {
Lingering(ctx context.Context, run system.Runner, name string) (bool, error)
SetLingering(ctx context.Context, run system.Runner, name string, on bool) error
}
// applyLinger makes whether an account lingers what was declared, read back from the machine, and
// keeps what it found the first time it changed it so removal can give that back.
func applyLinger(ctx context.Context, sys system.System, name string, want bool, run Runner,
out *Outcome) (bool, error) {
l, ok := sys.(lingerer)
if !ok {
return false, fmt.Errorf("%q is declared to linger, and this machine's service manager has no "+
"manager per account to keep running (novox/hq ADR 0177)", name)
}
now, err := l.Lingering(ctx, system.Runner(run), name)
if err != nil {
return false, err
}
if now == want {
return false, nil
}
if err := l.SetLingering(ctx, system.Runner(run), name, want); err != nil {
return false, err
}
if back, err := l.Lingering(ctx, system.Runner(run), name); err != nil {
return false, err
} else if back != want {
return false, fmt.Errorf("asked logind to make %q linger %s, and it reads %s",
name, onOff(want), onOff(back))
}
// Found once, as the shell's is: what goes back is what was there before the mesh.
if out.linger == nil {
out.linger = &store.Lingering{Found: now}
}
out.linger.Set = want
return true, nil
}
func onOff(on bool) string {
if on {
return "on"
}
return "off"
}
// removeUser is what undeclaring a login does: never deleting the account, and giving back the
// shell the mesh replaced when that is still safe (novox/hq ADR 0176 §2, issue 228) — and whether
// it lingered, on the same rule (novox/hq ADR 0177).
//
// **The account is never deleted, whether or not the mesh created it.** An account owns a home,
// files, a crontab, a mailbox — what a person did with it is not the mesh's to know, and deleting
// it is the data loss ADR 0030 exists to prevent. It is the package's rule, on a login: the
// mesh no longer requires it, which is not the same as "remove it".
//
// The shell goes back only while the account still has the one the mesh set — one a person chose
// since is theirs — and only to a shell that is still usable: giving back a shell that has been
// uninstalled since would break the very logins the giving back is for. Otherwise it is left, and
// the outcome says why. Never errNoRemoval: an orphaned login that failed removal stopped the
// whole apply, on every apply after.
func removeUser(ctx context.Context, sys system.System, a store.Applied, run Runner) (string, string, error) {
const kept = "the account is kept; the host never deletes a login"
login, exists, err := system.LookUpUser(ctx, system.Runner(run), a.Target)
if err != nil {
return "", "", err
}
if !exists {
return "forgotten", "no longer there", nil
}
action, detail, err := giveShellBack(ctx, sys, a, login, run, kept)
if err != nil {
return "", "", err
}
if gave, said := giveLingerBack(ctx, sys, a, run); said != "" {
detail += "; " + said
if gave {
action = "restored"
}
}
return action, detail, nil
}
// giveShellBack is removeUser's shell: given back only while the account still has the one the
// mesh set, and only to one still usable.
func giveShellBack(ctx context.Context, sys system.System, a store.Applied, login system.Login,
run Runner, kept string) (string, string, error) {
found := a.Shell
switch {
case found == nil:
return "forgotten", kept + ", and its shell was never changed by the mesh", nil
case found.Created:
return "forgotten", kept + "; the mesh created it, so there is no shell from before to give back", nil
case login.Shell != found.Set:
return "forgotten", fmt.Sprintf("%s, and its shell %s left as it is: changed since the mesh set %s",
kept, login.Shell, found.Set), nil
case found.Found == "":
return "forgotten", kept + ", and its shell left as it is: it had none before the mesh set one", nil
}
if err := system.UsableShell(found.Found); err != nil {
return "forgotten", fmt.Sprintf("%s, and its shell %s left as it is: the one it had before "+
"cannot be given back: %v", kept, login.Shell, err), nil
}
// A give-back that fails is said and not fatal: fatal, the record would stay and fail the same
// way on every apply after — the very wedge this removal exists to end.
if err := sys.SetUserShell(ctx, system.Runner(run), a.Target, found.Found); err != nil {
return "forgotten", fmt.Sprintf("%s, and the shell it had before the mesh, %s, could not be "+
"given back: %v", kept, found.Found, err), nil
}
if back, _, err := system.LookUpUser(ctx, system.Runner(run), a.Target); err != nil {
return "", "", err
} else if back.Shell != found.Found {
return "forgotten", fmt.Sprintf("%s; gave back the shell %s and the user database says %s",
kept, found.Found, back.Shell), nil
}
return "restored", fmt.Sprintf("%s; the shell it had before the mesh, %s, given back", kept, found.Found), nil
}
// giveLingerBack is removeUser's lingering (novox/hq ADR 0177), on the shell's rule: whether the
// account lingered before the mesh changed it goes back, and only while the account still has what
// the mesh set — one the operator changed since with loginctl is theirs. Never fatal, for the
// shell's reason. Empty when the mesh never changed it, so a removal says nothing about it.
//
// Giving back "not lingering" stops the account's manager if nobody is logged in, and every unit
// in it: that is what the account had before the mesh, and its user units go with its declaration.
func giveLingerBack(ctx context.Context, sys system.System, a store.Applied, run Runner) (bool, string) {
found := a.Linger
if found == nil {
return false, ""
}
if found.Found == found.Set {
return false, ""
}
l, ok := sys.(lingerer)
if !ok {
return false, ""
}
now, err := l.Lingering(ctx, system.Runner(run), a.Target)
if err != nil {
return false, fmt.Sprintf("whether it lingered before the mesh could not be given back: %v", err)
}
if now != found.Set {
return false, fmt.Sprintf("lingering left %s: changed since the mesh set it %s", onOff(now), onOff(found.Set))
}
if err := l.SetLingering(ctx, system.Runner(run), a.Target, found.Found); err != nil {
return false, fmt.Sprintf("lingering, %s before the mesh, could not be given back: %v", onOff(found.Found), err)
}
if back, err := l.Lingering(ctx, system.Runner(run), a.Target); err != nil || back != found.Found {
return false, fmt.Sprintf("gave back lingering %s and logind does not read it so", onOff(found.Found))
}
return true, fmt.Sprintf("lingering %s again, as before the mesh", onOff(found.Found))
}
// own sets a path's owner, when one was declared. // own sets a path's owner, when one was declared.
// //
// Looked up by name every time rather than cached: a user's numeric id is not stable across // Looked up by name every time rather than cached: a user's numeric id is not stable across
@@ -170,6 +375,75 @@ func ownedBy(path, owner string) (bool, error) {
return uid == wantUID && gid == wantGID, nil return uid == wantUID && gid == wantGID, nil
} }
// makeDirs makes a directory and any parent of it that is missing, as MkdirAll does — and gives
// each one it made inside the owner's home to the owner (novox/hq ADR 0182, to-be 41).
//
// **A parent made as root inside a home is a home the person cannot use.** A module writing
// ~/.config/mesh/environment.sh, or unpacking into ~/.local/share/powerlevel10k, on a fresh account
// made ~/.config and ~/.local/share owned by root: the file was the person's, the directory every
// program of theirs writes into was not. So what the host creates between the home and the target
// is the owner's, as the target is.
//
// **Only what the host created.** A parent that was already there is never chowned or chmodded:
// what a person or another program made is held as found (ADR 0182). And only inside the owner's
// home, read from the user database, not guessed from a prefix on /home: a module's directory under
// /var/lib is made exactly as before, whoever its files belong to.
func makeDirs(dir string, mode os.FileMode, owner string) error {
_, err := makeDirsSaying(dir, mode, owner)
return err
}
// makeDirsSaying is makeDirs, and says which directories it made, deepest first — so an archive
// can take away on removal the parents it made to reach its directory (novox/hq issue 162).
func makeDirsSaying(dir string, mode os.FileMode, owner string) ([]string, error) {
var made []string
for d := filepath.Clean(dir); ; d = filepath.Dir(d) {
if _, err := os.Lstat(d); !errors.Is(err, os.ErrNotExist) {
break
}
made = append(made, d)
if filepath.Dir(d) == d {
break
}
}
if err := os.MkdirAll(dir, mode); err != nil {
return nil, err
}
if owner == "" || len(made) == 0 {
return made, nil
}
home, err := homeOf(owner)
if err != nil || home == "" {
// A numeric owner — a container's user — has no home, and a name the machine does not
// know fails where the target is given to it. Either way nothing here is a home's.
return made, nil
}
home = filepath.Clean(home)
for _, d := range made {
if d != home && !strings.HasPrefix(d, home+string(os.PathSeparator)) {
continue
}
if err := ownMade(d, owner); err != nil {
return made, err
}
}
return made, nil
}
// homeOf is an owner's home from the user database, and ownMade gives a directory the host made to
// its owner. Variables so a test can give an owner a home it owns, and see what was given to whom
// without being root.
var (
homeOf = func(owner string) (string, error) {
found, err := osuser.Lookup(owner)
if err != nil {
return "", err
}
return found.HomeDir, nil
}
ownMade = own
)
// ownAll gives a whole tree to a user, for an archive that was unpacked into it. // ownAll gives a whole tree to a user, for an archive that was unpacked into it.
func ownAll(root, owner string) error { func ownAll(root, owner string) error {
if owner == "" { if owner == "" {
+294
View File
@@ -0,0 +1,294 @@
package apply
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
"github.com/novox/mesh-host/internal/system"
)
// Defends novox/hq ADR 0176 §2 and issue 228: a login the mesh set is given back when its holding
// moves, undeclaring one never stops the node applying, and a shell is checked before it is set.
// logins is a fake user database: each account's shell by name, and every command it was asked.
type logins struct {
shells map[string]string
asked []string
}
func (l *logins) run(_ context.Context, name string, args ...string) (string, error) {
l.asked = append(l.asked, name+" "+strings.Join(args, " "))
who := args[len(args)-1]
switch name {
case "getent":
if shell, ok := l.shells[who]; ok {
return who + ":x:1500:1500::/home/" + who + ":" + shell + "\n", nil
}
return "", errors.New("getent exited 2: ") // the host's runner's words for "no such key"
case "useradd":
shell := ""
for i, a := range args {
if a == "--shell" {
shell = args[i+1]
}
}
l.shells[who] = shell
case "usermod":
if args[0] == "--shell" {
l.shells[who] = args[1]
}
case "userdel":
delete(l.shells, who)
case "id":
return "\n", nil
}
return "", nil
}
func (l *logins) did(prefix string) bool {
for _, a := range l.asked {
if strings.HasPrefix(a, prefix) {
return true
}
}
return false
}
// shellsOn makes a machine's shells in a directory a test owns: each name an executable file,
// listed or not in the machine's list of shells as said, which this test's apply then reads.
func shellsOn(t *testing.T, listed []string, unlisted ...string) string {
t.Helper()
dir := t.TempDir()
var list strings.Builder
list.WriteString("# Pathnames of valid login shells.\n")
for _, name := range append(append([]string{}, listed...), unlisted...) {
if err := os.WriteFile(filepath.Join(dir, name), []byte("#!/bin/sh\n"), 0o755); err != nil {
t.Fatal(err)
}
}
for _, name := range listed {
list.WriteString(filepath.Join(dir, name) + "\n")
}
if err := os.WriteFile(filepath.Join(dir, "shells"), []byte(list.String()), 0o644); err != nil {
t.Fatal(err)
}
t.Cleanup(system.ShellsIn(filepath.Join(dir, "shells")))
return dir
}
func applyUsers(t *testing.T, l *logins, known store.State, resources string) (Report, store.State, error) {
t.Helper()
if resources == "" {
// Undeclared: something else stays, since a declaration with nothing in it is refused.
resources = `{"id":"other.dir","type":"directory","path":"` + t.TempDir() + `/other"}`
}
return Apply(context.Background(), archHost(t), parse(t, `{"declaration":1,"resources":[`+resources+`]}`),
known, store.OriginDeclared, l.run, nil, nil)
}
func userWith(shell string) string {
return `{"id":"shell.login","type":"user","name":"operator","shell":"` + shell + `"}`
}
func TestAnUndeclaredUserNoLongerStopsTheApply(t *testing.T) {
// Before issue 228 the host had no removal for a user, the orphan failed with "no way to
// remove", and an orphan's failure aborts the apply before its first resource — on every
// apply after, since the record stayed.
dir := shellsOn(t, []string{"bash", "zsh"})
l := &logins{shells: map[string]string{"operator": dir + "/bash"}}
_, state, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh"))
if err != nil {
t.Fatal(err)
}
page := filepath.Join(t.TempDir(), "page")
report, state, err := applyUsers(t, l, state,
`{"id":"web.page","type":"file","path":"`+page+`","content":"hello\n"}`)
if err != nil {
t.Fatalf("an undeclared user stopped the apply: %v", err)
}
if _, err := os.Stat(page); err != nil {
t.Errorf("a file in the same declaration was not written: %v", err)
}
if o := outcomeOf(report, "shell.login"); o.Action == "" {
t.Errorf("the user's removal was not reported: %+v", report.Outcomes)
}
if _, still := state.Find("shell.login"); still {
t.Error("the user is still recorded, so the next apply would meet it again")
}
}
func TestTheShellFoundIsGivenBackWhenTheUserIsUndeclared(t *testing.T) {
dir := shellsOn(t, []string{"bash", "zsh"})
l := &logins{shells: map[string]string{"operator": dir + "/bash"}}
_, state, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh"))
if err != nil {
t.Fatal(err)
}
if l.shells["operator"] != dir+"/zsh" {
t.Fatalf("the declared shell was not set: %q", l.shells["operator"])
}
if r, _ := state.Find("shell.login"); r.Shell == nil || r.Shell.Found != dir+"/bash" {
t.Fatalf("the shell the account had was not recorded: %+v", r.Shell)
}
report, _, err := applyUsers(t, l, state, "")
if err != nil {
t.Fatal(err)
}
if l.shells["operator"] != dir+"/bash" {
t.Errorf("the shell the account had was not given back: %q", l.shells["operator"])
}
if o := outcomeOf(report, "shell.login"); o.Action != "restored" {
t.Errorf("the give-back was not said: %+v", o)
}
if l.did("userdel") {
t.Error("the account was deleted")
}
}
func TestAShellAPersonChangedSinceIsLeftAlone(t *testing.T) {
dir := shellsOn(t, []string{"bash", "zsh", "fish"})
l := &logins{shells: map[string]string{"operator": dir + "/bash"}}
_, state, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh"))
if err != nil {
t.Fatal(err)
}
l.shells["operator"] = dir + "/fish" // chsh, by the person whose login it is
l.asked = nil
report, _, err := applyUsers(t, l, state, "")
if err != nil {
t.Fatal(err)
}
if l.did("usermod") || l.shells["operator"] != dir+"/fish" {
t.Errorf("a shell a person chose was taken from them: %q, %v", l.shells["operator"], l.asked)
}
if o := outcomeOf(report, "shell.login"); o.Action != "forgotten" || !strings.Contains(o.Detail, "changed since") {
t.Errorf("the outcome does not say why the shell was left: %+v", o)
}
}
func TestAFoundShellThatIsGoneIsNotGivenBack(t *testing.T) {
// Giving back a shell uninstalled since would break the logins the giving back is for.
dir := shellsOn(t, []string{"bash", "zsh"})
l := &logins{shells: map[string]string{"operator": dir + "/bash"}}
_, state, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh"))
if err != nil {
t.Fatal(err)
}
if err := os.Remove(dir + "/bash"); err != nil {
t.Fatal(err)
}
l.asked = nil
report, _, err := applyUsers(t, l, state, "")
if err != nil {
t.Fatal(err)
}
if l.did("usermod") || l.shells["operator"] != dir+"/zsh" {
t.Errorf("a shell no longer on the machine was given back: %q", l.shells["operator"])
}
if o := outcomeOf(report, "shell.login"); !strings.Contains(o.Detail, "cannot be given back") {
t.Errorf("the outcome does not say why the shell was left: %+v", o)
}
}
func TestAShellThatIsMissingOrUnlistedIsRefusedBeforeItIsSet(t *testing.T) {
dir := shellsOn(t, []string{"bash"}, "unlisted")
for name, shell := range map[string]string{
"missing": dir + "/zsh",
"unlisted": dir + "/unlisted",
} {
t.Run(name, func(t *testing.T) {
l := &logins{shells: map[string]string{"operator": dir + "/bash"}}
page := filepath.Join(t.TempDir(), "page")
report, state, err := applyUsers(t, l, store.State{}, userWith(shell)+`,
{"id":"web.page","type":"file","path":"`+page+`","content":"hello\n"}`)
if err == nil {
t.Fatal("the refused shell did not fail its resource")
}
if l.did("usermod") || l.shells["operator"] != dir+"/bash" {
t.Errorf("the account was changed: %q, %v", l.shells["operator"], l.asked)
}
if _, recorded := state.Find("shell.login"); recorded {
t.Error("a refused user was recorded")
}
if o := outcomeOf(report, "web.page"); o.Action != "created" {
t.Errorf("the refusal stopped the rest of the declaration: %+v", report.Outcomes)
}
})
}
t.Run("an account not yet made", func(t *testing.T) {
l := &logins{shells: map[string]string{}}
if _, _, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh")); err == nil {
t.Fatal("the missing shell was not refused")
}
if l.did("useradd") {
t.Errorf("the account was made with a shell that is not there: %v", l.asked)
}
})
}
func TestAnAccountThatRefusesLoginsNeedNotBeListed(t *testing.T) {
// A service's account has nologin, which no distribution lists among its shells; refusing it
// would refuse the controller's own account.
dir := shellsOn(t, []string{"bash"}, "nologin")
l := &logins{shells: map[string]string{}}
if _, _, err := applyUsers(t, l, store.State{}, userWith(dir+"/nologin")); err != nil {
t.Fatalf("a service account was refused: %v", err)
}
if l.shells["operator"] != dir+"/nologin" {
t.Errorf("the account was not made: %v", l.asked)
}
}
func TestACreatedAccountSurvivesItsRemoval(t *testing.T) {
dir := shellsOn(t, []string{"zsh"})
l := &logins{shells: map[string]string{}}
report, state, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh"))
if err != nil {
t.Fatal(err)
}
if o := outcomeOf(report, "shell.login"); o.Action != "created" {
t.Fatalf("the account was not created: %+v", o)
}
l.asked = nil
report, _, err = applyUsers(t, l, state, "")
if err != nil {
t.Fatal(err)
}
if _, still := l.shells["operator"]; !still || l.did("userdel") || l.did("usermod") {
t.Errorf("a created account was not left as it is: %v", l.asked)
}
if o := outcomeOf(report, "shell.login"); !strings.Contains(o.Detail, "account is kept") {
t.Errorf("the outcome does not say the account was kept: %+v", o)
}
}
func TestTheFoundShellIsNotOverwrittenByASecondChange(t *testing.T) {
// The holding moves from one shell module to another: what is given back in the end is the
// shell from before the mesh, not the first module's.
dir := shellsOn(t, []string{"bash", "zsh", "fish"})
l := &logins{shells: map[string]string{"operator": dir + "/bash"}}
_, state, err := applyUsers(t, l, store.State{}, userWith(dir+"/zsh"))
if err != nil {
t.Fatal(err)
}
_, state, err = applyUsers(t, l, state, userWith(dir+"/fish"))
if err != nil {
t.Fatal(err)
}
if r, _ := state.Find("shell.login"); r.Shell == nil || r.Shell.Found != dir+"/bash" || r.Shell.Set != dir+"/fish" {
t.Fatalf("the record is not the shell found and the one set last: %+v", r.Shell)
}
if _, _, err := applyUsers(t, l, state, ""); err != nil {
t.Fatal(err)
}
if l.shells["operator"] != dir+"/bash" {
t.Errorf("given back %q, not the shell from before the mesh", l.shells["operator"])
}
}
+558
View File
@@ -0,0 +1,558 @@
package apply
import (
"context"
"errors"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
"github.com/novox/mesh-host/internal/system"
)
// Defends novox/hq ADR 0177: a unit in the operator account's own service manager is applied
// through that manager — `systemctl --user --machine=<account>@` — and never as a system unit of
// the same name; its record remembers the scope so removal goes the same way. The account's
// manager runs only while somebody is logged in or the account lingers: with it away a unit waits
// rather than fails, a removal is never fatal, and the manager is never started by asking it.
// account is a machine with one account whose own manager runs or does not, and the units in it.
type account struct {
name, uid, home string
// up is whether the account's manager runs: a login, or lingering.
up bool
// lingers is logind's record, kept as files in a directory a test owns.
lingerDir string
// units are the account's units by name; system are the machine's.
units, system map[string]*fakeUnit
// busDown is a manager that says it runs while its bus does not answer — it stopped between
// the question and the command.
busDown bool
asked []string
// strays are system-scope commands that reached a unit, which no test here expects unless it
// declares a system unit.
strays []string
}
func newAccount(t *testing.T, up bool) *account {
t.Helper()
was := serviceSettle
serviceSettle = 0
dir := t.TempDir()
restore := system.LingerIn(dir)
t.Cleanup(func() { serviceSettle = was; restore() })
return &account{name: "ops", uid: "1001", home: "/home/ops", up: up, lingerDir: dir,
units: map[string]*fakeUnit{"i3-reload-watcher.service": {active: "inactive", enabled: "disabled"}},
system: map[string]*fakeUnit{}}
}
func (a *account) did(prefix string) bool {
for _, c := range a.asked {
if strings.HasPrefix(c, prefix) {
return true
}
}
return false
}
func (a *account) run(_ context.Context, name string, args ...string) (string, error) {
line := name + " " + strings.Join(args, " ")
a.asked = append(a.asked, line)
switch name {
case "getent":
if args[len(args)-1] == a.name {
return a.name + ":x:" + a.uid + ":" + a.uid + "::" + a.home + ":/bin/bash\n", nil
}
return "", errors.New("getent exited 2: ")
case "loginctl":
path := filepath.Join(a.lingerDir, args[1])
switch args[0] {
case "enable-linger":
a.up = true
return "", os.WriteFile(path, nil, 0o644)
case "disable-linger":
a.up = false
return "", os.Remove(path)
}
case "systemctl":
if line == "systemctl is-active user@"+a.uid+".service" {
if a.up {
return "active\n", nil
}
return "inactive\n", errors.New("systemctl exited 3: ")
}
prefix := "--machine=" + a.name + "@"
if len(args) > 1 && args[0] == "--user" && args[1] == prefix {
if !a.up || a.busDown {
return "", errors.New("systemctl exited 1: Failed to connect to user scope bus via " +
"machine transport: No such file or directory")
}
return unitCommand(a.units, args[2:])
}
a.strays = append(a.strays, line)
return unitCommand(a.system, args)
}
return "", nil
}
// unitCommand is one systemctl verb against a set of units.
func unitCommand(units map[string]*fakeUnit, args []string) (string, error) {
if len(args) == 0 {
return "", nil
}
if args[0] == "daemon-reload" {
return "", nil
}
u, ok := units[args[1]]
switch args[0] {
case "show":
if !ok {
return "LoadState=not-found\nActiveState=inactive", nil
}
return "LoadState=loaded\nActiveState=" + u.active, nil
case "is-enabled":
if !ok {
return "", errors.New("systemctl exited 1: ")
}
return u.enabled + "\n", nil
}
if !ok {
return "", fmt.Errorf("systemctl exited 5: Unit %s not found", args[1])
}
switch args[0] {
case "start", "restart":
u.active = "active"
case "stop":
u.active = "inactive"
case "enable":
u.enabled = "enabled"
case "disable":
u.enabled = "disabled"
}
return "", nil
}
const watcher = `{"id":"i3.watcher","type":"service","unit":"i3-reload-watcher.service","state":"running","boot":"enabled","scope":"user","user":"ops"}`
func declaring(resources ...string) string {
return `{"declaration":1,"resources":[` + strings.Join(resources, ",") + `]}`
}
func applyAccount(t *testing.T, raw string, known store.State, a *account) (Report, store.State, error) {
t.Helper()
return Apply(context.Background(), archHost(t), parse(t, raw), known, store.OriginDeclared, a.run, nil, nil)
}
func outcomeFor(r Report, id string) Outcome {
for _, o := range r.Outcomes {
if o.ID == id {
return o
}
}
return Outcome{}
}
func TestAUserScopedUnitIsAppliedThroughTheAccountsManager(t *testing.T) {
a := newAccount(t, true)
report, known, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err != nil {
t.Fatalf("apply: %v\n%s", err, strings.Join(a.asked, "\n"))
}
if !report.Changed() {
t.Fatal("a unit that was stopped and is now running changed nothing")
}
if !a.did("systemctl --user --machine=ops@ start i3-reload-watcher.service") ||
!a.did("systemctl --user --machine=ops@ enable i3-reload-watcher.service") {
t.Fatalf("the unit was not started and enabled in the account's manager:\n%s", strings.Join(a.asked, "\n"))
}
if len(a.strays) != 0 {
t.Fatalf("a user-scoped unit reached the machine's manager: %v", a.strays)
}
recorded, ok := known.At("service", "i3-reload-watcher.service")
if !ok || recorded.Scope != "user" || recorded.User != "ops" || recorded.Found == nil {
t.Fatalf("the record does not say whose manager the unit is in, or what was found: %+v", recorded)
}
}
func TestASystemUnitIsUntouchedByTheScope(t *testing.T) {
var commands []string
decl := `{"declaration":1,"resources":[
{"id":"x.daemon","type":"service","unit":"sshd.service","state":"running","boot":"enabled"}
]}`
if _, _, err := Apply(context.Background(), archHost(t), parse(t, decl),
store.State{}, store.OriginDeclared, unitIn(true, &commands), nil, nil); err != nil {
t.Fatal(err)
}
for _, c := range commands {
if strings.Contains(c, "--user") || strings.Contains(c, "--machine") || strings.Contains(c, "user@") {
t.Fatalf("a system unit was addressed to an account's manager: %s", c)
}
}
}
// With nobody logged in and no lingering, the unit waits: not a failure, nothing recorded, and the
// account's manager never asked — asking it would log the account in.
func TestAUserUnitWaitsForItsAccountsManagerAndIsAppliedWhenItRuns(t *testing.T) {
a := newAccount(t, false)
report, known, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err != nil {
t.Fatalf("a unit whose account is not logged in failed the apply: %v", err)
}
o := outcomeFor(report, "i3.watcher")
if o.Action != "waiting" || !strings.Contains(o.Detail, "not running") {
t.Fatalf("the outcome does not say the unit waits for its manager: %+v", o)
}
if report.Changed() {
t.Error("a unit that waited is reported as a change")
}
if a.did("systemctl --user") {
t.Fatalf("the account's manager was asked while it was not running:\n%s", strings.Join(a.asked, "\n"))
}
if _, ok := known.Find("i3.watcher"); ok {
t.Fatal("a unit never applied was recorded")
}
// The person logs in.
a.up = true
report, known, err = applyAccount(t, declaring(watcher), known, a)
if err != nil {
t.Fatal(err)
}
if o := outcomeFor(report, "i3.watcher"); o.Action == "waiting" || a.units["i3-reload-watcher.service"].active != "active" {
t.Fatalf("the unit was not applied once its manager ran: %+v", o)
}
if _, ok := known.Find("i3.watcher"); !ok {
t.Fatal("the unit was applied and not recorded")
}
}
// A unit applied before, its account since logged out: the record stays exactly as it was, what
// was found included, so a later removal still gives back what was there before the mesh.
func TestAWaitingUnitKeepsItsRecord(t *testing.T) {
a := newAccount(t, true)
_, known, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err != nil {
t.Fatal(err)
}
before, _ := known.Find("i3.watcher")
a.up = false
_, known, err = applyAccount(t, declaring(watcher), known, a)
if err != nil {
t.Fatal(err)
}
after, ok := known.Find("i3.watcher")
if !ok || after.Found == nil || *after.Found != *before.Found || after.Scope != "user" {
t.Fatalf("waiting changed the record: before %+v, after %+v", before, after)
}
}
// A manager that says it runs and then does not answer — the person logged out mid-apply — is the
// same absence, and is said the same way.
func TestAManagerThatStopsDuringTheApplyIsWaitedFor(t *testing.T) {
a := newAccount(t, true)
a.busDown = true
calls := 0
run := func(ctx context.Context, name string, args ...string) (string, error) {
if name == "systemctl" && len(args) == 2 && args[0] == "is-active" {
calls++
if calls > 1 {
a.up = false
}
}
return a.run(ctx, name, args...)
}
report, _, err := Apply(context.Background(), archHost(t), parse(t, declaring(watcher)), store.State{},
store.OriginDeclared, run, nil, nil)
if err != nil {
t.Fatalf("a manager that went away mid-apply failed it: %v", err)
}
if o := outcomeFor(report, "i3.watcher"); o.Action != "waiting" {
t.Fatalf("not waiting: %+v", o)
}
}
func TestAUserUnitOfAnAccountTheMachineDoesNotHaveIsRefused(t *testing.T) {
a := newAccount(t, true)
a.name = "someone-else"
_, _, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err == nil || !strings.Contains(err.Error(), `"ops"`) {
t.Fatalf("a unit of an account that is not here was not refused naming it: %v", err)
}
}
// Without a manager per account — OpenRC — a user-scoped unit would otherwise be applied as the
// machine's service of the same name. Refused instead.
func TestAUserUnitIsRefusedWhereTheServiceManagerHasNoAccounts(t *testing.T) {
alpine, err := system.For("alpine")
if err != nil {
t.Fatal(err)
}
a := newAccount(t, true)
_, _, err = Apply(context.Background(), alpine, parse(t, declaring(watcher)), store.State{},
store.OriginDeclared, a.run, nil, nil)
if err == nil || !strings.Contains(err.Error(), "no manager per account") {
t.Fatalf("not refused: %v", err)
}
if a.did("rc-service") {
t.Fatalf("a user-scoped unit reached the machine's services: %v", a.asked)
}
}
// **Removal is never fatal** (the issue 162/228 wedge): with the manager away the record stays and
// the outcome says it waits; the first apply that finds the manager gives the unit back.
func TestRemovingAUserUnitWithItsManagerAwayWaitsAndIsNeverFatal(t *testing.T) {
a := newAccount(t, true)
_, known, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err != nil {
t.Fatal(err)
}
a.up = false
report, known, err := applyAccount(t, nothingButA(t), known, a)
if err != nil {
t.Fatalf("undeclaring a unit whose account is logged out failed the apply: %v", err)
}
if o := outcomeFor(report, "i3.watcher"); o.Action != "waiting" || !strings.Contains(o.Detail, "not running") {
t.Fatalf("the removal does not say it waits: %+v", o)
}
if _, ok := known.Find("i3.watcher"); !ok {
t.Fatal("a unit not given back was forgotten")
}
a.up = true
report, known, err = applyAccount(t, nothingButA(t), known, a)
if err != nil {
t.Fatal(err)
}
u := a.units["i3-reload-watcher.service"]
if u.active != "inactive" || u.enabled != "disabled" {
t.Fatalf("the unit was not given back as found: %+v (%+v)", u, outcomeFor(report, "i3.watcher"))
}
if _, ok := known.Find("i3.watcher"); ok {
t.Fatal("a unit given back is still recorded")
}
if len(a.strays) != 0 {
t.Fatalf("the removal reached the machine's manager: %v", a.strays)
}
}
// "Failed to connect to bus" from a manager that said it ran is said and retried, never fatal.
func TestRemovingAUserUnitWhoseBusDoesNotAnswerIsNotFatal(t *testing.T) {
a := newAccount(t, true)
_, known, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err != nil {
t.Fatal(err)
}
a.busDown = true
report, known, err := applyAccount(t, nothingButA(t), known, a)
if err != nil {
t.Fatalf("a bus that did not answer made the removal fatal: %v", err)
}
if o := outcomeFor(report, "i3.watcher"); o.Action != "waiting" || !strings.Contains(o.Detail, "Failed to connect") {
t.Fatalf("the removal does not say what stopped it: %+v", o)
}
if _, ok := known.Find("i3.watcher"); !ok {
t.Fatal("a unit not given back was forgotten")
}
}
func TestRemovingAUserUnitOfAnAccountThatIsGoneForgetsIt(t *testing.T) {
a := newAccount(t, true)
_, known, err := applyAccount(t, declaring(watcher), store.State{}, a)
if err != nil {
t.Fatal(err)
}
a.name = "renamed"
report, known, err := applyAccount(t, nothingButA(t), known, a)
if err != nil {
t.Fatal(err)
}
if o := outcomeFor(report, "i3.watcher"); o.Action != "forgotten" || !strings.Contains(o.Detail, "no longer on this machine") {
t.Fatalf("%+v", o)
}
if _, ok := known.Find("i3.watcher"); ok {
t.Fatal("still recorded")
}
}
// A unit file the mesh wrote under the account's own unit directory makes the unit the mesh's — and
// only the account's unit of that name, never the machine's.
func TestAUserUnitsFileTheMeshWroteMakesThatUnitAndNoOtherTheMeshs(t *testing.T) {
home := t.TempDir()
was := homeOf
homeOf = func(name string) (string, error) {
if name == "ops" {
return home, nil
}
return "", errors.New("no such account")
}
t.Cleanup(func() { homeOf = was })
known := store.State{Resources: []store.Applied{
{ID: "f", Type: "file", Target: filepath.Join(home, ".config/systemd/user/watcher.service")},
{ID: "g", Type: "file", Target: "/etc/systemd/user/shared.service"},
{ID: "h", Type: "file", Target: filepath.Join(home, ".config/systemd/user/kept.service"), Kept: "/x"},
{ID: "s1", Type: "service", Target: "watcher.service", Scope: "user", User: "ops"},
{ID: "s2", Type: "service", Target: "shared.service", Scope: "user", User: "ops"},
{ID: "s3", Type: "service", Target: "kept.service", Scope: "user", User: "ops"},
{ID: "s4", Type: "service", Target: "watcher.service"},
}}
made := meshMadeUnits(known)
for key, want := range map[string]bool{
unitKey("user", "ops", "watcher.service"): true,
unitKey("user", "ops", "shared.service"): true,
unitKey("user", "ops", "kept.service"): false,
unitKey("", "", "watcher.service"): false,
unitKey("system", "", "shared.service"): false,
} {
if made[key] != want {
t.Errorf("%s: made %v, want %v", key, made[key], want)
}
}
if installedByHand(known, filepath.Join(home, ".config/systemd/user/watcher.service")) {
t.Error("a user unit file the mesh wrote reads as installed by hand")
}
if !installedByHand(known, filepath.Join(home, ".config/systemd/user/other.service")) {
t.Error("a user unit file somebody else put there reads as not installed by hand")
}
if installedByHand(known, "/usr/lib/systemd/user/pipewire.service") {
t.Error("a packaged user unit reads as installed by hand")
}
}
// Undeclared together, the account's unit whose file the mesh wrote is stopped and disabled in the
// account's manager — whatever was found — and a system unit of the same name is not touched.
func TestAMeshMadeUserUnitIsStoppedInItsAccountAndTheMachinesNamesakeIsNot(t *testing.T) {
a := newAccount(t, true)
home := t.TempDir()
was := homeOf
homeOf = func(string) (string, error) { return home, nil }
t.Cleanup(func() { homeOf = was })
unitFile := filepath.Join(home, ".config/systemd/user/i3-reload-watcher.service")
if err := os.MkdirAll(filepath.Dir(unitFile), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(unitFile, []byte("[Service]\n"), 0o644); err != nil {
t.Fatal(err)
}
a.units["i3-reload-watcher.service"] = &fakeUnit{active: "active", enabled: "enabled"}
a.system["i3-reload-watcher.service"] = &fakeUnit{active: "active", enabled: "enabled"}
known := store.State{Resources: []store.Applied{
{ID: "i3.unit", Type: "file", Target: unitFile, Origin: store.OriginDeclared},
{ID: "i3.watcher", Type: "service", Target: "i3-reload-watcher.service", Scope: "user", User: "ops",
Origin: store.OriginDeclared, Found: &store.FoundUnit{State: "running", Boot: "enabled"}},
}}
if _, _, err := applyAccount(t, nothingButA(t), known, a); err != nil {
t.Fatal(err)
}
if u := a.units["i3-reload-watcher.service"]; u.active != "inactive" || u.enabled != "disabled" {
t.Fatalf("the mesh's own user unit was not stopped and disabled: %+v", u)
}
if u := a.system["i3-reload-watcher.service"]; u.active != "active" || u.enabled != "enabled" {
t.Fatalf("the machine's unit of the same name was touched: %+v %v", u, a.strays)
}
}
// A service moved from the machine's manager into an account's is two units: the machine's is given
// back as it was found, through the machine's manager, and the account's is read afresh.
func TestAServiceMovedIntoAnAccountGivesTheMachinesUnitBack(t *testing.T) {
a := newAccount(t, true)
a.system["i3-reload-watcher.service"] = &fakeUnit{active: "active", enabled: "enabled"}
known := store.State{Resources: []store.Applied{
{ID: "i3.watcher", Type: "service", Target: "i3-reload-watcher.service", Origin: store.OriginDeclared,
Found: &store.FoundUnit{Unit: "i3-reload-watcher.service", State: "stopped", Boot: "disabled"}},
}}
_, known, err := applyAccount(t, declaring(watcher), known, a)
if err != nil {
t.Fatal(err)
}
if u := a.system["i3-reload-watcher.service"]; u.active != "inactive" || u.enabled != "disabled" {
t.Fatalf("the machine's unit was not given back as found: %+v", u)
}
if u := a.units["i3-reload-watcher.service"]; u.active != "active" {
t.Fatalf("the account's unit was not started: %+v", u)
}
r, _ := known.Find("i3.watcher")
if r.Found == nil || r.Found.State != "stopped" || r.Scope != "user" {
t.Fatalf("what was found is not the account's unit's: %+v", r)
}
}
// Lingering is the account's (novox/hq ADR 0177): declared, set and read back; recorded with what
// was found; given back on removal while the account still has what the mesh set.
func TestLingeringIsDeclaredOnTheAccountAndGivenBack(t *testing.T) {
a := newAccount(t, false)
user := `{"id":"ops.login","type":"user","name":"ops","linger":true}`
report, known, err := applyAccount(t, declaring(user), store.State{}, a)
if err != nil {
t.Fatal(err)
}
if !a.did("loginctl enable-linger ops") || outcomeFor(report, "ops.login").Action != "updated" {
t.Fatalf("lingering was not enabled: %v", a.asked)
}
r, _ := known.Find("ops.login")
if r.Linger == nil || r.Linger.Found || !r.Linger.Set {
t.Fatalf("the record does not say what was found and set: %+v", r.Linger)
}
a.asked = nil
report, known, err = applyAccount(t, declaring(user), known, a)
if err != nil {
t.Fatal(err)
}
if a.did("loginctl") || outcomeFor(report, "ops.login").Action != "unchanged" {
t.Fatalf("a second apply changed lingering: %v", a.asked)
}
// And a user-scoped unit of the account now applies with nobody logged in.
if _, known, err = applyAccount(t, declaring(user, watcher), known, a); err != nil {
t.Fatal(err)
}
if a.units["i3-reload-watcher.service"].active != "active" {
t.Fatal("a lingering account's unit was not started")
}
report, _, err = applyAccount(t, nothingButA(t), known, a)
if err != nil {
t.Fatal(err)
}
if !a.did("loginctl disable-linger ops") {
t.Fatalf("lingering was not given back: %v", a.asked)
}
if o := outcomeFor(report, "ops.login"); o.Action != "restored" || !strings.Contains(o.Detail, "lingering off") {
t.Fatalf("%+v", o)
}
}
func TestLingeringTheOperatorChangedSinceIsLeft(t *testing.T) {
a := newAccount(t, false)
known := store.State{Resources: []store.Applied{{ID: "ops.login", Type: "user", Target: "ops",
Origin: store.OriginDeclared, Linger: &store.Lingering{Found: false, Set: true}}}}
// Not lingering now: somebody ran disable-linger since the mesh set it.
report, _, err := applyAccount(t, nothingButA(t), known, a)
if err != nil {
t.Fatal(err)
}
if a.did("loginctl") {
t.Fatalf("lingering the operator changed was changed back: %v", a.asked)
}
if o := outcomeFor(report, "ops.login"); !strings.Contains(o.Detail, "changed since") {
t.Fatalf("%+v", o)
}
}
func TestLingeringIsRefusedWhereTheServiceManagerHasNoAccounts(t *testing.T) {
alpine, err := system.For("alpine")
if err != nil {
t.Fatal(err)
}
a := newAccount(t, false)
_, _, err = Apply(context.Background(), alpine, parse(t, declaring(
`{"id":"ops.login","type":"user","name":"ops","linger":true}`)), store.State{},
store.OriginDeclared, a.run, nil, nil)
if err == nil || !strings.Contains(err.Error(), "linger") {
t.Fatalf("not refused: %v", err)
}
}
+113
View File
@@ -0,0 +1,113 @@
package apply
import (
"errors"
"fmt"
"os"
"strconv"
"github.com/novox/mesh-host/internal/store"
)
// A file written whole, undeclared (novox/hq ADR 0118, ADR 0102).
//
// **What the mesh made goes; what it wrote over is given back.** Before the host writes a file over
// one it has no record of making, it keeps the original first (ADR 0102: "whatever the host writes
// over without a record of it, it keeps first"). Undeclaring gives a thing back the state it was
// found in (ADR 0118), so a file with a kept original is not deleted when its record goes: the
// original is put back, with the mode and owner it was found with. Deleting it was the failure —
// a module that writes the package manager's configuration whole, unassigned, left the machine with
// no configuration at all.
//
// The cases, decided once and in this order:
//
// - **No kept original** — the mesh made the file where there was none (or the record is from
// before the host kept originals, which it cannot tell apart): removed, as before.
// - **The file is gone** — somebody removed it: nothing is put back, since bringing back a file a
// person deleted is not giving back the state the mesh found; the original stays kept.
// - **The file was changed since the mesh last wrote it** — it is somebody's again, as a block or
// a JSON file the mesh wrote into stays somebody's: left exactly as it stands, never clobbered,
// and the outcome names where the original is so a person can choose.
// - **The kept copy cannot be read** — the mesh's file is left in place rather than deleted, and
// the outcome says the original is missing.
// - Otherwise the original is written back atomically, and the outcome is "restored".
//
// **Never fatal.** Each case that leaves the file says so and lets the record go; none stops the
// rest of an unassignment. The kept copy itself is never deleted (novox/hq ADR 0100).
func removeWhole(a store.Applied) (string, string, error) {
if a.Kept == "" {
if err := os.RemoveAll(a.Target); err != nil {
return "", "", err
}
if _, err := os.Stat(a.Target); !errors.Is(err, os.ErrNotExist) {
return "", "", fmt.Errorf("%s is still there after removing it", a.Target)
}
return "removed", "no longer declared", nil
}
current, err := os.ReadFile(a.Target)
if errors.Is(err, os.ErrNotExist) {
return "forgotten", "no longer there; the original the mesh wrote over stays kept at " + a.Kept, nil
}
if err != nil {
return "kept", fmt.Sprintf("no longer declared, and it cannot be read (%v), so it was left as it "+
"is; the original the mesh wrote over is kept at %s", err, a.Kept), nil
}
if a.Wrote == "" || digestOf(string(current)) != a.Wrote {
return "kept", "no longer declared, and changed on the machine since the mesh last wrote it, so " +
"it was left as it is; the original the mesh wrote over is kept at " + a.Kept, nil
}
original, err := os.ReadFile(a.Kept)
if err != nil {
return "kept", fmt.Sprintf("no longer declared, but the original it was written over cannot be "+
"read at %s (%v), so the mesh's file was left in place", a.Kept, err), nil
}
info, err := os.Stat(a.Target)
if err != nil {
return "kept", fmt.Sprintf("no longer declared, and it cannot be seen (%v), so it was left as it "+
"is; the original the mesh wrote over is kept at %s", err, a.Kept), nil
}
mode := info.Mode().Perm()
if a.KeptMode != "" {
if m, err := strconv.ParseUint(a.KeptMode, 8, 32); err == nil {
mode = os.FileMode(m).Perm()
}
}
if err := writeAtomically(a.Target, original, mode); err != nil {
return "kept", fmt.Sprintf("no longer declared, and the original kept at %s could not be put "+
"back (%v), so the mesh's file was left in place", a.Kept, err), nil
}
detail := "no longer declared; the original the mesh wrote over was put back from " + a.Kept
if err := giveOwnerBack(a.Target, a.KeptOwner, info); err != nil {
detail += "; " + err.Error()
}
if back, err := os.ReadFile(a.Target); err != nil || string(back) != string(original) {
return "kept", "no longer declared; putting back the original kept at " + a.Kept +
" did not leave it there — check the file by hand", nil
}
return "restored", detail, nil
}
// giveOwnerBack gives a file put back the owner its original was found with — "uid:gid" as a hold
// records it — or, on a record from before the host kept that, the owner of what it replaced.
func giveOwnerBack(path, owner string, was os.FileInfo) error {
if owner == "" {
return keepOwner(path, was)
}
uid, gid, err := idsOf(owner)
if err != nil {
return fmt.Errorf("its owner %q could not be read: %w", owner, err)
}
now, err := os.Stat(path)
if err != nil {
return err
}
if u, g, ok := ownerOf(now); ok && u == uid && g == gid {
return nil
}
if err := os.Chown(path, uid, gid); err != nil {
return fmt.Errorf("its owner %s could not be given back: %w", owner, err)
}
return nil
}
+173
View File
@@ -0,0 +1,173 @@
package apply
import (
"os"
"path/filepath"
"strings"
"testing"
"github.com/novox/mesh-host/internal/store"
)
// Defends novox/hq ADR 0118 with ADR 0102: a file the host wrote whole over one it found is given
// its kept original back when it is undeclared — not deleted, which left a machine whose package
// manager's configuration a module wrote with no configuration at all once that module was
// unassigned.
const pacmanFound = "[options]\nArchitecture = auto\n\n[core]\nInclude = /etc/pacman.d/mirrorlist\n"
const pacmanMesh = "# written by the mesh\n[options]\nArchitecture = auto\nParallelDownloads = 5\n"
// writtenOver is a file found at path with content and mode, then written whole by the mesh.
func writtenOver(t *testing.T, content string, mode os.FileMode) (path string, state store.State) {
t.Helper()
path = filepath.Join(t.TempDir(), "pacman.conf")
if err := os.WriteFile(path, []byte(content), mode); err != nil {
t.Fatal(err)
}
if err := os.Chmod(path, mode); err != nil {
t.Fatal(err)
}
_, state = applyKeepingIn(t, wholeDecl(path, pacmanMesh), store.State{}, t.TempDir())
if got := readText(t, path); got != pacmanMesh {
t.Fatalf("the mesh's file was not written: %q", got)
}
return path, state
}
func TestAFileWrittenOverGetsItsKeptOriginalBackWhenUndeclared(t *testing.T) {
path, state := writtenOver(t, pacmanFound, 0o640)
rec, _ := state.Find(namesID)
if rec.Kept == "" || rec.KeptMode != "0640" {
t.Fatalf("the original and how it was found were not recorded: kept %q, mode %q", rec.Kept, rec.KeptMode)
}
if steps := Plan(somethingElse(t), state, store.OriginDeclared); !strings.Contains(verbs(steps), "restore "+namesID) {
t.Errorf("the plan did not say the original goes back: %s", verbs(steps))
}
report, after := undeclare(t, state)
if got := readText(t, path); got != pacmanFound {
t.Fatalf("undeclared, the machine did not get its original back: %q", got)
}
info, err := os.Stat(path)
if err != nil || info.Mode().Perm() != 0o640 {
t.Errorf("the original came back with mode %o, it was found 640", info.Mode().Perm())
}
o := outcomeOf(report, namesID)
if o.Action != "restored" || !strings.Contains(o.Detail, rec.Kept) {
t.Errorf("the give-back was reported as %q: %s", o.Action, o.Detail)
}
if _, still := after.Find(namesID); still {
t.Error("the record outlived its declaration")
}
if _, err := os.Stat(rec.Kept); err != nil {
t.Errorf("the kept copy went with the give-back: %v", err)
}
}
func TestAFileTheMeshMadeIsRemovedWhenUndeclared(t *testing.T) {
path := filepath.Join(t.TempDir(), "pacman.conf")
_, state := applyKeepingIn(t, wholeDecl(path, pacmanMesh), store.State{}, t.TempDir())
if rec, _ := state.Find(namesID); rec.Kept != "" {
t.Fatalf("a file that was not there recorded an original at %s", rec.Kept)
}
report, _ := undeclare(t, state)
if _, err := os.Stat(path); !os.IsNotExist(err) {
t.Errorf("a file the mesh made outlived its declaration: %v", err)
}
if got := outcomeOf(report, namesID).Action; got != "removed" {
t.Errorf("removal was reported as %q", got)
}
}
func TestAFileWrittenOverAndChangedSinceIsLeftAsItStands(t *testing.T) {
path, state := writtenOver(t, pacmanFound, 0o644)
edited := pacmanMesh + "IgnorePkg = linux\n"
if err := os.WriteFile(path, []byte(edited), 0o644); err != nil {
t.Fatal(err)
}
report, _ := undeclare(t, state)
if got := readText(t, path); got != edited {
t.Fatalf("the operator's change was clobbered: %q", got)
}
rec, _ := state.Find(namesID)
o := outcomeOf(report, namesID)
if o.Action != "kept" || !strings.Contains(o.Detail, "changed on the machine") || !strings.Contains(o.Detail, rec.Kept) {
t.Errorf("leaving it was reported as %q: %s", o.Action, o.Detail)
}
}
func TestAFileWhoseKeptOriginalIsMissingIsLeftAndSaysSo(t *testing.T) {
path, state := writtenOver(t, pacmanFound, 0o644)
rec, _ := state.Find(namesID)
if err := os.Remove(rec.Kept); err != nil {
t.Fatal(err)
}
report, _ := undeclare(t, state)
if got := readText(t, path); got != pacmanMesh {
t.Fatalf("with no original to put back, the file became %q", got)
}
o := outcomeOf(report, namesID)
if o.Action != "kept" || !strings.Contains(o.Detail, "cannot be read at "+rec.Kept) {
t.Errorf("leaving it was reported as %q: %s", o.Action, o.Detail)
}
}
func TestAFileWrittenOverAndDeletedSinceIsNotBroughtBack(t *testing.T) {
path, state := writtenOver(t, pacmanFound, 0o644)
if err := os.Remove(path); err != nil {
t.Fatal(err)
}
report, _ := undeclare(t, state)
if _, err := os.Stat(path); !os.IsNotExist(err) {
t.Errorf("a file somebody deleted was brought back: %v", err)
}
if got := outcomeOf(report, namesID).Action; got != "forgotten" {
t.Errorf("reported as %q", got)
}
}
func TestAFileWhosePathMovedIsNeverGivenTheOldPathsOriginal(t *testing.T) {
// The old path's original stays with the old path's record; the new path keeps its own.
oldPath, state := writtenOver(t, pacmanFound, 0o644)
newPath := filepath.Join(filepath.Dir(oldPath), "pacman.d.conf")
newFound := "# the machine's own at the new path\n"
if err := os.WriteFile(newPath, []byte(newFound), 0o644); err != nil {
t.Fatal(err)
}
_, state = applyKeepingIn(t, wholeDecl(newPath, pacmanMesh), state, t.TempDir())
rec, _ := state.Find(namesID)
if rec.Kept == "" {
t.Fatal("the original at the new path was written over without being kept")
}
if kept := readText(t, rec.Kept); kept != newFound {
t.Fatalf("the new path's record names the wrong original: %q", kept)
}
// The old path is a former target, given back by the next apply; the new one by undeclaring.
undeclare(t, state)
if got := readText(t, newPath); got != newFound {
t.Errorf("undeclared, the new path holds %q", got)
}
if got := readText(t, oldPath); got != pacmanFound {
t.Errorf("the old path did not get its own original back: %q", got)
}
}
func TestARecordFromBeforeTheModeWasKeptPutsTheOriginalBackAsTheFileStands(t *testing.T) {
path, state := writtenOver(t, pacmanFound, 0o644)
for i := range state.Resources {
state.Resources[i].KeptMode, state.Resources[i].KeptOwner = "", ""
}
if err := os.Chmod(path, 0o600); err != nil {
t.Fatal(err)
}
report, _ := undeclare(t, state)
if got := readText(t, path); got != pacmanFound {
t.Fatalf("got %q", got)
}
info, _ := os.Stat(path)
if info.Mode().Perm() != 0o600 {
t.Errorf("mode %o, the file stood at 600", info.Mode().Perm())
}
if got := outcomeOf(report, namesID).Action; got != "restored" {
t.Errorf("reported as %q", got)
}
}
+39
View File
@@ -6,6 +6,8 @@ import (
"encoding/json" "encoding/json"
"errors" "errors"
"fmt" "fmt"
"os"
"path/filepath"
"strings" "strings"
) )
@@ -112,6 +114,43 @@ func BuildControlPlane(ctx context.Context, run Runner, builderTag string, sourc
return Built{}, nil return Built{}, nil
} }
// **Which form the controller is in at that commit** (novox/hq issue 223). An image the module
// builds is the form genesis always raised, and the builder builds it as before. A process the
// module runs from a Go bundle cannot be built here — its toolchain is one the mesh makes later —
// so genesis builds the controller's own Dockerfile and raises it as the container that process
// replaces (genesis_form.go).
workspace, err := os.MkdirTemp("", "mesh-genesis-*")
if err != nil {
return Built{}, err
}
defer os.RemoveAll(workspace)
dir, commit, err := cloneAt(ctx, run, builderTag, source, workspace)
if err != nil {
return Built{}, fmt.Errorf("the control plane could not be fetched from %s at %s: %w",
source.Repository, shortRef(source.Ref), err)
}
raw, err := os.ReadFile(filepath.Join(dir, "module.json"))
if err != nil {
return Built{}, fmt.Errorf("%s at %s has no module manifest: %w", source.Repository, shortRef(source.Ref), err)
}
form, err := processFormOf(raw)
if err != nil {
return Built{}, err
}
if form.Found {
image, err := buildGenesisImage(ctx, run, dir)
if err != nil {
return Built{}, err
}
manifest, err := genesisForm(raw, form, image)
if err != nil {
return Built{}, err
}
say(fmt.Sprintf(" built %s from %s, as the container its process %s replaces (%s)",
ControlPlaneModule, shortRef(commit), form.Process, form.Replaces))
return Built{Module: ControlPlaneModule, Commit: commit, Image: image, Manifest: manifest}, nil
}
out, err := run(ctx, "docker", args...) out, err := run(ctx, "docker", args...)
if err != nil { if err != nil {
return Built{}, fmt.Errorf("the control plane could not be built from %s at %s: %w", return Built{}, fmt.Errorf("the control plane could not be built from %s at %s: %w",
+6 -6
View File
@@ -49,10 +49,10 @@ func TestARepositoryAndACommitIsEnough(t *testing.T) {
func TestTheBuildHandsOverTheManifestTheMeshWillHold(t *testing.T) { func TestTheBuildHandsOverTheManifestTheMeshWillHold(t *testing.T) {
manifest := `{"module":"mesh-controller","version":"1","resources":[` + manifest := `{"module":"mesh-controller","version":"1","resources":[` +
`{"id":"server","type":"container","name":"mesh-controller","image":"` + builtImage + `"}]}` `{"id":"server","type":"container","name":"mesh-controller","image":"` + builtImage + `"}]}`
runtime := &asked{answer: func(string, []string) (string, error) { runtime := &asked{answer: aRepository(t, imageFormManifest, func(string, []string) (string, error) {
return `{"module":"mesh-controller","commit":"a1b2c3d4","manifest":` + manifest + return `{"module":"mesh-controller","commit":"a1b2c3d4","manifest":` + manifest +
`,"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}` + "\n", nil `,"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}` + "\n", nil
}} })}
built, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test", built, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test",
Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {}) Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {})
if err != nil { if err != nil {
@@ -69,10 +69,10 @@ func TestTheBuildHandsOverTheManifestTheMeshWillHold(t *testing.T) {
// A result without a manifest is a build the installer cannot finish, and it is refused beside the // A result without a manifest is a build the installer cannot finish, and it is refused beside the
// builder that said it rather than at step 9 with a message about a missing file. // builder that said it rather than at step 9 with a message about a missing file.
func TestABuildReportingNoManifestIsRefused(t *testing.T) { func TestABuildReportingNoManifestIsRefused(t *testing.T) {
runtime := &asked{answer: func(string, []string) (string, error) { runtime := &asked{answer: aRepository(t, imageFormManifest, func(string, []string) (string, error) {
return `{"module":"mesh-controller","commit":"a1b2c3d4",` + return `{"module":"mesh-controller","commit":"a1b2c3d4",` +
`"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}`, nil `"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}`, nil
}} })}
_, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test", _, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test",
Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {}) Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {})
if err == nil || !strings.Contains(err.Error(), "no manifest") { if err == nil || !strings.Contains(err.Error(), "no manifest") {
@@ -82,11 +82,11 @@ func TestABuildReportingNoManifestIsRefused(t *testing.T) {
// And a manifest that does not name the image the build produced describes some other build. // And a manifest that does not name the image the build produced describes some other build.
func TestABuildWhoseManifestNamesAnotherImageIsRefused(t *testing.T) { func TestABuildWhoseManifestNamesAnotherImageIsRefused(t *testing.T) {
runtime := &asked{answer: func(string, []string) (string, error) { runtime := &asked{answer: aRepository(t, imageFormManifest, func(string, []string) (string, error) {
return `{"module":"mesh-controller","commit":"a1b2c3d4","manifest":{"module":"mesh-controller",` + return `{"module":"mesh-controller","commit":"a1b2c3d4","manifest":{"module":"mesh-controller",` +
`"resources":[{"id":"server","type":"container","image":"sha256:` + strings.Repeat("9", 64) + `"}]},` + `"resources":[{"id":"server","type":"container","image":"sha256:` + strings.Repeat("9", 64) + `"}]},` +
`"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}`, nil `"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}`, nil
}} })}
_, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test", _, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test",
Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {}) Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {})
if err == nil || !strings.Contains(err.Error(), "does not name that image") { if err == nil || !strings.Contains(err.Error(), "does not name that image") {
+205
View File
@@ -0,0 +1,205 @@
package bootstrap
import (
"bytes"
"context"
"encoding/json"
"fmt"
"os"
"path/filepath"
"sort"
"strings"
)
// The controller as a process, raised at genesis as a container (novox/hq issue 213, issue 223).
//
// **The mesh runs the controller as a Go bundle the host starts as a process; genesis cannot.** A
// process's bundle is fetched from the mesh's artifact store, which genesis raises long after the
// controller, and a Go bundle is compiled in a toolchain the mesh builds later still. So genesis
// pivots to the controller as it always has — an image it built from the controller's own
// Dockerfile, run as a container the temporary controller composes — and the first time the mesh
// builds the controller from its repository, the controller's own declaration is the process, which
// names that container under `replaces`, and the host hands over: the process is started, seen up,
// and only then is the container removed (mesh-host `replaces`, issue 213).
//
// **The container is genesis's shape, not the manifest's.** The manifest declares only the process.
// From it genesis takes what the process is given — its environment, which is host paths and words —
// and the id the process replaces; the container around it is written here: the image genesis built,
// the host's network, and every host path the environment names mounted at the same path read-only.
// It runs as the image's own unprivileged user (65534), so the secrets belong to that number until
// the process's account takes them over — the image is FROM scratch and knows no account by name. Its
// id is the one the process replaces, so the first declaration the mesh composes for this machine
// hands this container over rather than leaving two controllers running.
// genesisUser is who the genesis container runs as — the image's own USER — and who its secrets
// belong to until the process takes them over: the image has no passwd to look an account up in.
const genesisUser = "65534:65534"
// ProcessForm is the controller's process, as its manifest declares it, and the container id that
// process replaces. Found is false for a manifest in the image form — an older controller — which
// genesis installs as it always did.
type ProcessForm struct {
Found bool
Process string // the process resource's id
Replaces string // the id of the container genesis raises in its place
}
// processFormOf finds the controller's process in its manifest: a process resource running a bundle
// the module builds, saying which one resource it replaces.
func processFormOf(manifest []byte) (ProcessForm, error) {
var m struct {
Build *struct {
Artifacts []struct {
Name string `json:"name"`
Kind string `json:"kind"`
} `json:"artifacts"`
} `json:"build"`
Resources []map[string]any `json:"resources"`
}
if err := json.Unmarshal(manifest, &m); err != nil {
return ProcessForm{}, fmt.Errorf("the %s module's manifest is not readable: %w", ControlPlaneModule, err)
}
kinds := map[string]string{}
if m.Build != nil {
for _, a := range m.Build.Artifacts {
kinds[a.Name] = a.Kind
}
}
for _, kind := range kinds {
if kind == "image" {
return ProcessForm{}, nil // the image form: the builder builds it, as before
}
}
var found []ProcessForm
for _, r := range m.Resources {
if r["type"] != "process" || kinds[fmt.Sprint(r["artifact"])] != "bundle" {
continue
}
if once, _ := r["run-once"].(bool); once {
continue
}
replaces, _ := r["replaces"].([]any)
if len(replaces) != 1 {
return ProcessForm{}, fmt.Errorf(
"the %s module runs as the process %v and says it replaces %v. Genesis raises the "+
"controller as a container that process takes over, so the process names exactly "+
"one resource it replaces — the id genesis gives the container",
ControlPlaneModule, r["id"], r["replaces"])
}
found = append(found, ProcessForm{Found: true, Process: fmt.Sprint(r["id"]),
Replaces: fmt.Sprint(replaces[0])})
}
if len(found) != 1 {
return ProcessForm{}, fmt.Errorf(
"the %s module builds no image and runs %d process(es) of its own; genesis raises one "+
"controller, from the process its manifest declares", ControlPlaneModule, len(found))
}
return found[0], nil
}
// genesisForm is the manifest genesis registers: the controller's own manifest, its process
// replaced by the container genesis runs in its place, under the id the process replaces, running
// the image genesis built. Resolved as a build would resolve it — no build section, the image named
// — because that is what the temporary controller is handed.
func genesisForm(manifest []byte, form ProcessForm, image string) ([]byte, error) {
var m map[string]any
if err := json.Unmarshal(manifest, &m); err != nil {
return nil, err
}
resources, _ := m["resources"].([]any)
var out []any
for _, raw := range resources {
r, _ := raw.(map[string]any)
if r == nil || r["id"] != form.Process {
out = append(out, raw)
continue
}
env, _ := r["env"].(map[string]any)
var volumes []any
for _, key := range sortedAnyKeys(env) {
value := fmt.Sprint(env[key])
if strings.HasPrefix(value, "/") || strings.HasPrefix(value, "${dir:") {
volumes = append(volumes, value+":"+value+":ro")
}
}
container := map[string]any{
"id": form.Replaces, "type": "container", "name": ControlPlaneModule,
"image": image, "network": "host", "args": []any{"serve"},
}
if len(env) > 0 {
container["env"] = env
}
if len(volumes) > 0 {
container["volumes"] = volumes
}
out = append(out, container)
}
m["resources"] = out
// What the container reads must be readable by who it runs as. The process's account owns them
// once the process takes over, and the host gives them to it in the same apply.
m["secrets-owner"] = genesisUser
// **Nothing to prepare at genesis.** The temporary controller — the same commit — migrated the
// stores when the foundation raised it, and a preparation step is derived from a resource running
// an artifact the module built, which a pinned image is not: the controller refuses `prepares`
// with nothing to run it in. The process prepares the stores itself when it takes over.
delete(m, "prepares")
delete(m, "build")
var b bytes.Buffer
enc := json.NewEncoder(&b)
enc.SetEscapeHTML(false)
if err := enc.Encode(m); err != nil {
return nil, err
}
return bytes.TrimSpace(b.Bytes()), nil
}
func sortedAnyKeys(m map[string]any) []string {
keys := make([]string, 0, len(m))
for k := range m {
keys = append(keys, k)
}
sort.Strings(keys)
return keys
}
// cloneAt fetches the controller's repository at the commit genesis builds, into a directory of
// this machine's, with the carried builder's git — the machine is not assumed to have one. Returns
// the module's directory and the commit that was checked out.
func cloneAt(ctx context.Context, run Runner, builderTag string, source Source, into string) (string, string, error) {
git := func(args ...string) (string, error) {
return run(ctx, "docker", append([]string{"run", "--rm", "-v", into + ":/ws",
"--entrypoint", "git", builderTag}, args...)...)
}
if _, err := git("clone", "--quiet", source.Repository, "/ws/src"); err != nil {
return "", "", fmt.Errorf("cloning %s: %w", source.Repository, err)
}
if _, err := git("-C", "/ws/src", "checkout", "--quiet", "--detach", source.Ref); err != nil {
return "", "", fmt.Errorf("checking out %s: %w", shortRef(source.Ref), err)
}
commit, err := git("-C", "/ws/src", "rev-parse", "HEAD")
if err != nil {
return "", "", err
}
return filepath.Join(into, "src", source.Path), strings.TrimSpace(commit), nil
}
// buildGenesisImage builds the controller's image from its own Dockerfile (the one `make image`
// uses), on this machine, and names it by the digest of its own configuration, as the builder did.
func buildGenesisImage(ctx context.Context, run Runner, dir string) (string, error) {
if _, err := os.Stat(filepath.Join(dir, "Dockerfile")); err != nil {
return "", fmt.Errorf("the %s repository has no Dockerfile, so genesis has no image to raise "+
"the controller from: %w", ControlPlaneModule, err)
}
out, err := run(ctx, "docker", "build", "--quiet", dir)
if err != nil {
return "", fmt.Errorf("building the %s image: %w", ControlPlaneModule, err)
}
image := strings.TrimSpace(out)
if i := strings.LastIndex(image, "\n"); i >= 0 {
image = strings.TrimSpace(image[i+1:])
}
if !strings.HasPrefix(image, "sha256:") {
return "", fmt.Errorf("docker build said %q, which is not an image id", firstLine(out))
}
return image, nil
}
+182
View File
@@ -0,0 +1,182 @@
package bootstrap
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
// imageFormManifest is a controller from before issue 213: it builds an image and runs a container.
const imageFormManifest = `{"module":"mesh-controller","version":"1","resources":[
{"id":"server","type":"container","name":"mesh-controller","network":"host","artifact":"server"}],
"build":{"artifacts":[{"name":"server","kind":"image","from":"Dockerfile"}]}}`
// processFormManifest is the controller's manifest as novox/mesh-controller declares it after issue
// 213: a Go bundle the host runs as a process, replacing the container it ran as.
const processFormManifest = `{
"module": "mesh-controller", "version": "1", "slug": "control", "prepares": true,
"claims": [{"name": "mesh-controller", "scope": "mesh"}],
"accesses": [{"path": "/var/lib/mesh-broker-tls", "mode": "read"}],
"own-secrets": {"inventory": "${dir:mesh-state}/inventory", "bus": "${dir:mesh-state}/bus"},
"secrets-owner": "mesh-controller",
"tools": ["status"],
"resources": [
{"id": "account", "type": "user", "name": "mesh-controller", "shell": "/usr/bin/nologin", "home": "/var/lib/mesh-controller"},
{"id": "mesh-state", "type": "directory", "mode": "0700", "place": "mesh", "owner": "mesh-controller"},
{"id": "controller", "type": "process", "name": "mesh-controller", "artifact": "controller",
"run": ["./mesh-controller", "serve"], "user": "mesh-controller",
"env": {"MESH_BROKER_CERTIFICATE": "/var/lib/mesh-broker-tls/tls.crt",
"MESH_STORE_INVENTORY_FILE": "${dir:mesh-state}/inventory",
"MESH_STORE_INVENTORY_PORT": "${seat:mesh-store:5432}",
"MESH_BUS_NATS_FILE": "${dir:mesh-state}/bus"},
"replaces": ["server"]}
],
"build": {"artifacts": [{"name": "controller", "kind": "bundle", "language": "go", "system": "arch",
"from": "cmd/mesh-controller", "binary": "mesh-controller"}]}
}`
// aRepository answers the carried builder's git as a clone of a repository holding this manifest and
// a Dockerfile, and hands every other command to then.
func aRepository(t *testing.T, manifest string, then func(string, []string) (string, error)) func(string, []string) (string, error) {
t.Helper()
return func(name string, args []string) (string, error) {
if name == "docker" && len(args) > 5 && args[0] == "run" && args[4] == "--entrypoint" && args[5] == "git" {
host := strings.TrimSuffix(args[3], ":/ws")
joined := strings.Join(args, " ")
switch {
case strings.Contains(joined, " clone "):
src := filepath.Join(host, "src")
if err := os.MkdirAll(src, 0o755); err != nil {
return "", err
}
if err := os.WriteFile(filepath.Join(src, "module.json"), []byte(manifest), 0o644); err != nil {
return "", err
}
return "", os.WriteFile(filepath.Join(src, "Dockerfile"), []byte("FROM scratch\n"), 0o644)
case strings.Contains(joined, "rev-parse"):
return "a1b2c3d4e5f6\n", nil
}
return "", nil
}
return then(name, args)
}
}
// novox/hq issue 223: a controller declared as a process is raised at genesis as a container built
// from its own Dockerfile, not by the builder — which has no Go toolchain at genesis and would refuse.
func TestAProcessFormControllerIsBuiltFromItsDockerfileAndRaisedAsAContainer(t *testing.T) {
runtime := &asked{answer: aRepository(t, processFormManifest, func(name string, args []string) (string, error) {
if name == "docker" && args[0] == "build" {
return builtImage + "\n", nil
}
t.Fatalf("genesis ran %s %v; a process-form controller is built from its Dockerfile alone", name, args)
return "", nil
})}
built, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test",
Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {})
if err != nil {
t.Fatal(err)
}
if built.Image != builtImage || built.Commit != "a1b2c3d4e5f6" {
t.Errorf("built %q from %q", built.Image, built.Commit)
}
if runtime.ran("mesh-builder:test build") {
t.Error("the builder was asked to build a controller it cannot build at genesis")
}
var m map[string]any
if err := json.Unmarshal(built.Manifest, &m); err != nil {
t.Fatalf("the genesis manifest is not JSON: %v", err)
}
if _, has := m["prepares"]; has {
t.Error("the genesis manifest prepares its state; the temporary controller already did, and a pinned image is nothing the controller can derive a step from")
}
if _, has := m["build"]; has {
t.Error("the genesis manifest still says how it is built; it is handed over resolved")
}
var container map[string]any
for _, raw := range m["resources"].([]any) {
r := raw.(map[string]any)
if r["type"] == "process" {
t.Errorf("the genesis manifest still runs the process: %v", r)
}
if r["type"] == "container" {
container = r
}
}
if container == nil {
t.Fatal("the genesis manifest runs no container")
}
for key, want := range map[string]any{"id": "server", "name": "mesh-controller", "image": builtImage,
"network": "host"} {
if container[key] != want {
t.Errorf("the genesis container's %s is %v, not %v", key, container[key], want)
}
}
if _, has := container["user"]; has {
t.Error("the genesis container says a user; the host's container has no such field, the image's USER is who it runs as")
}
if m["secrets-owner"] != "65534:65534" {
t.Errorf("the secrets belong to %v, which the container cannot read as", m["secrets-owner"])
}
volumes, _ := json.Marshal(container["volumes"])
for _, want := range []string{"${dir:mesh-state}/inventory:${dir:mesh-state}/inventory:ro",
"/var/lib/mesh-broker-tls/tls.crt:/var/lib/mesh-broker-tls/tls.crt:ro"} {
if !strings.Contains(string(volumes), want) {
t.Errorf("the container does not mount %s: %s", want, volumes)
}
}
if strings.Contains(string(volumes), "seat:") {
t.Errorf("a word that is not a path was mounted: %s", volumes)
}
// And it is what step 9 installs: pinned, its container found, its stores delivered from its
// environment — the same path an image-form controller takes.
pinned, _, err := pinImage(built.Manifest, built.Image, "registry.internal:5000/mesh-controller@sha256:"+strings.Repeat("e", 64), ControlPlaneModule)
if err != nil {
t.Fatal(err)
}
if id := controlPlaneResourceIn(pinned); id != "server" {
t.Errorf("step 9 finds the controller's container as %q", id)
}
wanted, err := secretsByVariableIn(pinned)
if err != nil {
t.Fatalf("step 9 cannot deliver the stores into the genesis container: %v", err)
}
if wanted["MESH_STORE_INVENTORY"] != "inventory" {
t.Errorf("step 9 delivers %v", wanted)
}
}
// An older controller — an image and a container — is still built by the builder, as before.
func TestAnImageFormControllerIsStillBuiltByTheBuilder(t *testing.T) {
manifest := `{"module":"mesh-controller","version":"1","resources":[` +
`{"id":"server","type":"container","name":"mesh-controller","image":"` + builtImage + `"}]}`
runtime := &asked{answer: aRepository(t, imageFormManifest, func(name string, args []string) (string, error) {
if name == "docker" && args[0] == "build" {
t.Fatal("an image-form controller was built from its Dockerfile rather than by the builder")
}
return `{"module":"mesh-controller","commit":"a1b2c3d4","manifest":` + manifest +
`,"made":[{"name":"server","kind":"image","reference":"` + builtImage + `"}]}`, nil
})}
built, err := BuildControlPlane(context.Background(), runtime.run, "mesh-builder:test",
Source{Repository: "https://example.invalid/mesh-controller.git", Ref: "a1b2c3d4"}, false, func(string) {})
if err != nil {
t.Fatal(err)
}
if string(built.Manifest) != manifest || !runtime.ran("mesh-builder:test build") {
t.Errorf("the image form did not go through the builder: %s", built.Manifest)
}
}
func TestAProcessFormNamingNoOneReplacementIsRefused(t *testing.T) {
for _, bad := range []string{`"replaces": []`, `"replaces": ["a", "b"]`} {
raw := strings.Replace(processFormManifest, `"replaces": ["server"]`, bad, 1)
if _, err := processFormOf([]byte(raw)); err == nil {
t.Errorf("a process saying %s was accepted; genesis would not know what to name its container", bad)
}
}
}
+121
View File
@@ -0,0 +1,121 @@
package bootstrap
import (
"archive/tar"
"bytes"
"compress/gzip"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/novox/mesh-host/internal/apply"
"github.com/novox/mesh-host/internal/declaration"
"github.com/novox/mesh-host/internal/store"
"github.com/novox/mesh-host/internal/system"
)
// novox/hq issue 223: what genesis raises is what the controller's process takes over. The temporary
// controller composes the genesis container, and the host records it as `<module>.<its id>`; the
// first declaration the mesh composes from the controller's real manifest names that same id under
// the process's `replaces` (the composer prefixes both alike — mesh-controller's
// TestTheControllerIsAProcessAndNoContainer). So the first apply hands over: the process is started,
// seen up, and only then is the genesis container removed — one controller before, one after, never
// none and never two left.
func TestTheFirstApplyHandsTheGenesisContainerOverToTheProcess(t *testing.T) {
restore := apply.ForTests(t.TempDir(), t.TempDir())
defer restore()
form, err := processFormOf([]byte(processFormManifest))
if err != nil {
t.Fatal(err)
}
genesis, err := genesisForm([]byte(processFormManifest), form, builtImage)
if err != nil {
t.Fatal(err)
}
// What the host records for the genesis container, as the composer names it.
recorded := ""
var m struct {
Resources []map[string]any `json:"resources"`
}
if err := json.Unmarshal(genesis, &m); err != nil {
t.Fatal(err)
}
for _, r := range m.Resources {
if r["type"] == "container" {
recorded = ControlPlaneModule + "." + r["id"].(string)
}
}
known := store.State{Resources: []store.Applied{{Origin: store.OriginDeclared, ID: recorded,
Type: "container", Target: ControlPlaneModule}}}
// The controller's first composed declaration: its process, replacing what the manifest names.
body, digest := aBundle(t)
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { _, _ = w.Write(body) }))
defer server.Close()
replaces, _ := json.Marshal([]string{ControlPlaneModule + "." + form.Replaces})
d, err := declaration.Parse([]byte(`{"declaration":1,"resources":[{"id":"` + ControlPlaneModule + "." + form.Process +
`","type":"process","name":"mesh-controller","source":"` + server.URL + `/c.tgz","digest":"` + digest +
`","run":["./mesh-controller","serve"],"replaces":` + string(replaces) + `}]}`))
if err != nil {
t.Fatal(err)
}
var commands []string
started, container := false, true
run := func(_ context.Context, name string, args ...string) (string, error) {
line := name + " " + strings.Join(args, " ")
commands = append(commands, line)
switch {
case strings.HasPrefix(line, "systemctl restart mesh-controller.service"):
started = true
case strings.HasPrefix(line, "systemctl show mesh-controller.service") && started:
return "ActiveState=active\nSubState=running\nMainPID=7\nNRestarts=0\n", nil
case strings.HasPrefix(line, "docker rm -f mesh-controller"):
if !started {
t.Error("the genesis container was removed before the process was started")
}
container = false
case strings.HasPrefix(line, "docker container inspect") && !container:
return "", errors.New("no such container")
}
return "", nil
}
sys, err := system.For("arch")
if err != nil {
t.Fatal(err)
}
_, after, err := apply.Apply(context.Background(), sys, d, known, store.OriginDeclared, run, nil, nil)
if err != nil {
t.Fatalf("the first apply did not hand over: %v\n%s", err, strings.Join(commands, "\n"))
}
if container {
t.Fatalf("the genesis container is still running beside the process — two controllers:\n%s",
strings.Join(commands, "\n"))
}
if _, still := after.Find(recorded); still {
t.Error("the host still records the genesis container")
}
}
func aBundle(t *testing.T) ([]byte, string) {
t.Helper()
var raw bytes.Buffer
zipped := gzip.NewWriter(&raw)
w := tar.NewWriter(zipped)
content := "#!/bin/sh\n"
if err := w.WriteHeader(&tar.Header{Name: "mesh-controller", Mode: 0o755, Size: int64(len(content)), Typeflag: tar.TypeReg}); err != nil {
t.Fatal(err)
}
_, _ = w.Write([]byte(content))
_ = w.Close()
_ = zipped.Close()
sum := sha256.Sum256(raw.Bytes())
return raw.Bytes(), "sha256:" + hex.EncodeToString(sum[:])
}
+203
View File
@@ -374,6 +374,15 @@ type User struct {
// Home directory. Absent means the system's default for a new user, and is not changed for // Home directory. Absent means the system's default for a new user, and is not changed for
// one that exists — moving somebody's home is not something a declaration should do quietly. // one that exists — moving somebody's home is not something a declaration should do quietly.
Home string `json:"home,omitempty"` Home string `json:"home,omitempty"`
// Linger is whether the account's own service manager runs with nobody logged in (novox/hq ADR
// 0177). A user-scoped unit lives in that manager, and the manager runs only from the account's
// first login to its last logout — so a server's user unit, where nobody ever logs in, never
// runs without it, and a workstation's runs only while its person is there, which on a desktop
// is what is wanted. Absent asserts nothing, as Shell's does: true has the account linger, false
// has it not. On the account rather than on the unit, because it is the account's: two units of
// one account cannot disagree about it, and undeclaring one of them must not stop the other.
Linger *bool `json:"linger,omitempty"`
} }
// Network is a named network on this machine. // Network is a named network on this machine.
@@ -569,6 +578,18 @@ type Process struct {
// that runs once does not run on a schedule, and something that is not running cannot be // that runs once does not run on a schedule, and something that is not running cannot be
// restarted when a file changes. // restarted when a file changes.
Schedule string `json:"schedule,omitempty"` Schedule string `json:"schedule,omitempty"`
// Replaces names resources this declaration no longer declares that this process takes the
// place of (novox/hq issue 213). Such a resource is not removed with the other orphans, before
// anything is applied: it is removed only once this process is applied and still running a
// moment later, and kept when it is not. So the thing being replaced answers until the thing
// replacing it does — the controller moving from its container to a process is the case: the
// container removed first left nothing answering the mesh's verbs for as long as fetching,
// unpacking and starting the process took, and for ever if the process did not start.
//
// For a process that stays up; a step or a scheduled run is not running a moment later by
// design, so there is nothing to hand over to.
Replaces []string `json:"replaces,omitempty"`
} }
func (d *Process) Identity() string { return d.ID } func (d *Process) Identity() string { return d.ID }
@@ -654,6 +675,20 @@ func (d *Process) validate(where string, _ bool) []string {
problems = append(problems, where+": "+err.Error()) problems = append(problems, where+": "+err.Error())
} }
} }
if len(d.Replaces) > 0 && (d.RunOnce || d.Schedule != "") {
problems = append(problems, where+
": only a process that stays up replaces something — a step or a scheduled run is not "+
"running a moment later, so what it replaced would be removed with nothing in its place, "+
"or never")
}
for _, id := range d.Replaces {
switch {
case strings.TrimSpace(id) == "":
problems = append(problems, where+": replaces names an empty id")
case id == d.ID:
problems = append(problems, where+": a process cannot replace itself")
}
}
return problems return problems
} }
@@ -684,6 +719,16 @@ type Service struct {
// declaration that reports success and stops being true at the next power cut. // declaration that reports success and stops being true at the next power cut.
Boot string `json:"boot,omitempty"` Boot string `json:"boot,omitempty"`
// Scope is whose service manager the unit belongs to: "system" (absent means system), or
// "user" — the operator account's own manager (novox/hq ADR 0177). A workstation's per-user
// daemons — a window manager's reload watcher, an audio mask, a memory guard — are units in
// that manager, and until this they had no form the mesh could send. A user-scoped unit names
// its User; the host talks to that account's manager and never starts a system unit by the
// same name.
Scope string `json:"scope,omitempty"`
// User is the account whose manager a user-scoped unit lives in. Required with scope "user",
// refused otherwise; a module names it ${machine:account} and never the person.
User string `json:"user,omitempty"`
// RestartOn names resources whose change means this service must be restarted. // RestartOn names resources whose change means this service must be restarted.
// //
// Because a running service does not re-read its configuration. Replace the file, find the // Because a running service does not re-read its configuration. Replace the file, find the
@@ -729,11 +774,33 @@ func (s *Service) Identity() string { return s.ID }
func (s *Service) Kind() Type { return TypeService } func (s *Service) Kind() Type { return TypeService }
func (s *Service) Target() string { return s.Unit } func (s *Service) Target() string { return s.Unit }
// ScopeSystem and ScopeUser are the two managers a unit may belong to (novox/hq ADR 0177).
const (
ScopeSystem = "system"
ScopeUser = "user"
)
// UserScoped is whether this unit lives in an account's own service manager.
func (s *Service) UserScoped() bool { return s.Scope == ScopeUser }
func (s *Service) validate(where string, _ bool) []string { func (s *Service) validate(where string, _ bool) []string {
var problems []string var problems []string
if s.Unit == "" { if s.Unit == "" {
problems = append(problems, where+": a service needs a unit") problems = append(problems, where+": a service needs a unit")
} }
switch s.Scope {
case "", ScopeSystem:
if s.User != "" {
problems = append(problems, where+": a system unit names no user; only a user-scoped unit does")
}
case ScopeUser:
if s.User == "" {
problems = append(problems, where+": a user-scoped unit names the account whose manager it lives in")
}
default:
problems = append(problems, fmt.Sprintf(
"%s: scope %q; a unit is in the \"system\" manager or the operator account's \"user\" one", where, s.Scope))
}
switch { switch {
case s.State == "running" || s.State == "stopped": case s.State == "running" || s.State == "stopped":
case s.State != "": case s.State != "":
@@ -870,6 +937,12 @@ type Package struct {
ID string `json:"id"` ID string `json:"id"`
Type Type `json:"type"` Type Type `json:"type"`
Package string `json:"package"` Package string `json:"package"`
// Absent declares that the package is NOT installed (novox/hq ADR 0180): the host removes it
// when it is, and leaves a machine that never had it alone. For the one case a module replaces
// software the machine was found with and the operator has decided it does not come back — the
// firewall front end a converged machine's filter module retired. Nothing to undo when the
// declaration drops it: the host does not install what a declaration stopped saying is absent.
Absent bool `json:"absent,omitempty"`
} }
func (p *Package) Identity() string { return p.ID } func (p *Package) Identity() string { return p.ID }
@@ -940,6 +1013,20 @@ type Container struct {
// its siblings can name before any of them can resolve anything. // its siblings can name before any of them can resolve anything.
Dns []string `json:"dns,omitempty"` Dns []string `json:"dns,omitempty"`
// Capabilities are the Linux capabilities this container is granted beyond the runtime's
// default set, by name (novox/hq ADR 0170): a holder's runtime that changes the machine's packet
// filter asks for NET_ADMIN. Exactly these, named in the spec so a change recreates the
// container; a privileged container stays undeclarable.
Capabilities []string `json:"capabilities,omitempty"`
// Logging names where the runtime sends this container's output: "journald" sends it to the
// machine's journal, under the container's name, where what reads the machine's logs — its
// intrusion prevention first of all (novox/hq ADR 0179) — can read it the way it reads the
// machine's own services. Empty keeps the runtime's default, which is a file of the runtime's
// own that nothing but the runtime reads. Part of the spec: a container that logs elsewhere
// is a different container, and the runtime cannot change a running one's driver.
Logging string `json:"logging,omitempty"`
// Networks are networks this container also joins once created, by name — a found network a // Networks are networks this container also joins once created, by name — a found network a
// per-machine setting keeps for a taken container (novox/hq ADR 0163, rule 4), so a // per-machine setting keeps for a taken container (novox/hq ADR 0163, rule 4), so a
// neighbour that resolves it there keeps resolving it until the neighbour is taken too. // neighbour that resolves it there keeps resolving it until the neighbour is taken too.
@@ -991,6 +1078,24 @@ type Container struct {
// rather than stacked. It is exclusive with RunOnce and with restart-on: a container runs once // rather than stacked. It is exclusive with RunOnce and with restart-on: a container runs once
// and gates, runs on a cadence, or stays up — never two of these. // and gates, runs on a cadence, or stays up — never two of these.
Schedule string `json:"schedule,omitempty"` Schedule string `json:"schedule,omitempty"`
// WhileStopped names resources of the same module — containers — that must be held still for
// the duration of this step's run (novox/hq ADR 0189). The host stops each before the run and
// starts each again after it, **whatever the step did**: a step that failed must leave the
// service running, because the one real risk of this field is a window that never closes.
//
// **For the work a service cannot have done underneath it.** The artifact store's collector
// walks the storage and requires every writer stopped; a run-once step runs beside containers
// and a scheduled one is the same container again, so until this there was no way for a module
// to say it. The predecessor said it with a shell script, which is how the mesh inherited a
// store that has never collected anything.
//
// **Its own module's containers, and only on a schedule.** A module that could quiesce a
// neighbour could stop the mesh. And at apply time the host already has a window — the
// declaration is applied in order and a run-once step gates what follows — so a one-time
// offline job says *before*, not *instead of*; a recurring window is the case order cannot
// express, and the only one this serves.
WhileStopped []string `json:"while-stopped,omitempty"`
} }
func (c *Container) Identity() string { return c.ID } func (c *Container) Identity() string { return c.ID }
@@ -1022,6 +1127,20 @@ func (c *Container) validate(where string, _ bool) []string {
problems = append(problems, where+": "+err.Error()) problems = append(problems, where+": "+err.Error())
} }
} }
// A maintenance window belongs to a recurring step (novox/hq ADR 0189). Refused on anything
// else here, where the field is; that it names containers of the same module, and not itself,
// is judged against the whole declaration (see whileStoppedNames).
if len(c.WhileStopped) > 0 && c.Schedule == "" {
problems = append(problems, where+": while-stopped needs a schedule; at apply the host "+
"already has a window — the declaration is applied in order and a run-once step gates "+
"what follows — so a one-time offline job is declared before what it works on")
}
for _, id := range c.WhileStopped {
if id == c.ID {
problems = append(problems, where+": while-stopped names "+strconv.Quote(id)+
", which is this step itself")
}
}
// The runtime's flags take addresses, and a name here would be handed to it verbatim and // The runtime's flags take addresses, and a name here would be handed to it verbatim and
// refused at create — after the old container was already removed. Refused on arrival instead. // refused at create — after the old container was already removed. Refused on arrival instead.
for _, d := range c.Dns { for _, d := range c.Dns {
@@ -1039,6 +1158,16 @@ func (c *Container) validate(where string, _ bool) []string {
"static address anywhere but a user-defined one") "static address anywhere but a user-defined one")
} }
} }
for _, cap := range c.Capabilities {
if !capabilityName.MatchString(cap) {
problems = append(problems, where+": capabilities names "+strconv.Quote(cap)+", which is not a "+
"capability's name (CAP_NET_ADMIN or NET_ADMIN)")
}
}
if c.Logging != "" && c.Logging != "journald" {
problems = append(problems, where+": logging is "+strconv.Quote(c.Logging)+", and the only place a "+
"container's output can be sent besides the runtime's own file is \"journald\"")
}
for _, n := range c.Networks { for _, n := range c.Networks {
problems = append(problems, (&Network{Name: n}).validate(where+": networks", false)...) problems = append(problems, (&Network{Name: n}).validate(where+": networks", false)...)
if n == c.Network { if n == c.Network {
@@ -1209,6 +1338,10 @@ type Adoption struct {
Untaken map[string][]string `json:"untaken,omitempty"` Untaken map[string][]string `json:"untaken,omitempty"`
} }
// capabilityName is what a Linux capability is called: upper case, underscores, an optional CAP_
// prefix. The runtime accepts either spelling.
var capabilityName = regexp.MustCompile(`^(CAP_)?[A-Z][A-Z0-9_]*$`)
// AdoptionPrefix is the id prefix of what the mesh itself declares because a node is adopted — // AdoptionPrefix is the id prefix of what the mesh itself declares because a node is adopted —
// its openings and its guard. Nothing under it belongs to a module, so none of it is ever held. // its openings and its guard. Nothing under it belongs to a module, so none of it is ever held.
const AdoptionPrefix = "adoption." const AdoptionPrefix = "adoption."
@@ -1422,7 +1555,9 @@ func parse(raw []byte, allowActions bool) (*Declaration, error) {
problems = append(problems, resource.validate(where, allowActions)...) problems = append(problems, resource.validate(where, allowActions)...)
d.Resources = append(d.Resources, resource) d.Resources = append(d.Resources, resource)
} }
problems = append(problems, checkWhileStopped(d.Resources)...)
problems = append(problems, checkAdoption(env.Adoption, d.Resources, allowActions)...) problems = append(problems, checkAdoption(env.Adoption, d.Resources, allowActions)...)
problems = append(problems, checkReplaces(d.Resources)...)
if env.Adoption == nil { if env.Adoption == nil {
for _, r := range d.Resources { for _, r := range d.Resources {
if r.Kind() == TypeOpening { if r.Kind() == TypeOpening {
@@ -1450,6 +1585,41 @@ func parse(raw []byte, allowActions bool) (*Declaration, error) {
return d, nil return d, nil
} }
// checkWhileStopped judges a maintenance window against the whole declaration (novox/hq ADR 0189).
//
// A step may hold still only a container that is **here** — in this same declaration, which is to
// say on this machine and placed by the mesh. That is what makes it the module's own: a node's
// declaration carries one module's resources beside another's, so the id must also be a container
// and not a file or a directory, which there would be nothing to stop.
//
// Refused on arrival rather than discovered at the first fire. A window that names something the
// host cannot stop is a window that opens at 03:00 and reports nothing until somebody reads a log.
func checkWhileStopped(resources []Resource) []string {
containers := map[string]bool{}
for _, r := range resources {
if r.Kind() == TypeContainer {
containers[r.Identity()] = true
}
}
var problems []string
for _, r := range resources {
c, ok := r.(*Container)
if !ok {
continue
}
for _, id := range c.WhileStopped {
if containers[id] {
continue
}
problems = append(problems, fmt.Sprintf(
"resource %q: while-stopped names %q, and this declaration has no container by "+
"that id. A step may hold still only a container placed on this machine "+
"beside it", c.ID, id))
}
}
return problems
}
func strictDecode(raw []byte, into any) error { func strictDecode(raw []byte, into any) error {
// DisallowUnknownFields is the whole point rather than strictness for its own sake: a // DisallowUnknownFields is the whole point rather than strictness for its own sake: a
// field the host does not know is a thing the control plane believes it asked for. // field the host does not know is a thing the control plane believes it asked for.
@@ -1606,3 +1776,36 @@ func stripComments(raw []byte) []byte {
} }
return []byte(strings.Join(kept, "\n")) return []byte(strings.Join(kept, "\n"))
} }
// checkReplaces refuses a `replaces` naming something this same declaration still declares, or
// named by two processes (novox/hq issue 213). What is replaced is what the declaration no longer
// says — a resource still declared is applied, not handed over, and one handed to two replacements
// would go when the first of them came up, whatever became of the second.
func checkReplaces(resources []Resource) []string {
declared := map[string]bool{}
for _, r := range resources {
declared[r.Identity()] = true
}
var problems []string
by := map[string]string{}
for _, r := range resources {
p, ok := r.(*Process)
if !ok {
continue
}
for _, id := range p.Replaces {
if declared[id] {
problems = append(problems, fmt.Sprintf(
"resource %q: replaces %q, which this declaration still declares — a process "+
"replaces what the declaration no longer says", p.ID, id))
}
if other, taken := by[id]; taken && other != p.ID {
problems = append(problems, fmt.Sprintf(
"resource %q: replaces %q, which %q replaces too — one thing has one replacement",
p.ID, id, other))
}
by[id] = p.ID
}
}
return problems
}
+41
View File
@@ -504,3 +504,44 @@ func TestKeptNetworksAndLeftOutModulesAreReadStrictly(t *testing.T) {
t.Fatalf("a carried bundle leaving modules out was accepted: %v", err) t.Fatalf("a carried bundle leaving modules out was accepted: %v", err)
} }
} }
// A container may ask for a capability by name, and nothing else (novox/hq ADR 0170).
func TestACapabilityIsNamedOrRefused(t *testing.T) {
image := "postgres@sha256:" + strings.Repeat("a", 64)
d, err := Parse([]byte(`{"declaration":1,"resources":[
{"id":"fw","type":"container","name":"fw","image":"` + image + `","network":"host","capabilities":["NET_ADMIN","CAP_NET_RAW"]}
]}`))
if err != nil {
t.Fatal(err)
}
if got := d.Resources[0].(*Container).Capabilities; len(got) != 2 || got[0] != "NET_ADMIN" {
t.Fatalf("capabilities read as %v", got)
}
for _, bad := range []string{`"net_admin"`, `"ALL;rm -rf /"`, `"privileged"`} {
if _, err := Parse([]byte(`{"declaration":1,"resources":[
{"id":"fw","type":"container","name":"fw","image":"` + image + `","capabilities":[` + bad + `]}]}`)); err == nil {
t.Errorf("%s was accepted as a capability", bad)
}
}
}
// A container may send its output to the machine's journal, and nowhere else but the runtime's own
// file (novox/hq ADR 0179): what reads the machine's logs then reads the container's too.
func TestAContainerMayLogToTheJournalAndNowhereElse(t *testing.T) {
image := "postgres@sha256:" + strings.Repeat("a", 64)
d, err := Parse([]byte(`{"declaration":1,"resources":[
{"id":"front","type":"container","name":"front","image":"` + image + `","logging":"journald"}
]}`))
if err != nil {
t.Fatal(err)
}
if got := d.Resources[0].(*Container).Logging; got != "journald" {
t.Fatalf("logging read as %q", got)
}
for _, bad := range []string{`"syslog"`, `"none"`, `"json-file"`} {
if _, err := Parse([]byte(`{"declaration":1,"resources":[
{"id":"front","type":"container","name":"front","image":"` + image + `","logging":` + bad + `}]}`)); err == nil {
t.Errorf("%s was accepted as a place to log", bad)
}
}
}
+63
View File
@@ -0,0 +1,63 @@
package declaration
import (
"strings"
"testing"
)
// novox/hq issue 213: a process says what it takes the place of, so the host can keep the old one
// answering until the new one does. What it may name is narrow, and each refusal is said here.
const aReplacingProcess = `{"id":"m.controller","type":"process","name":"m","source":"https://store.invalid/m",
"digest":"sha256:` + "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa" + `","run":["./m","serve"]%s}`
func replacing(extra string) string {
return `{"declaration":1,"resources":[` + strings.Replace(aReplacingProcess, "%s", extra, 1) + `]}`
}
func TestAProcessMayNameWhatItReplaces(t *testing.T) {
d, err := Parse([]byte(replacing(`,"replaces":["m.server"]`)))
if err != nil {
t.Fatalf("a process replacing an undeclared container was refused: %v", err)
}
if p := d.Resources[0].(*Process); len(p.Replaces) != 1 || p.Replaces[0] != "m.server" {
t.Fatalf("replaces was not read: %+v", p)
}
}
// What is replaced is what the declaration no longer says. A resource still declared is applied,
// and handing it over as well would remove something the same declaration asks to keep.
func TestAProcessMayNotReplaceSomethingStillDeclared(t *testing.T) {
raw := `{"declaration":1,"resources":[
{"id":"m.server","type":"directory","path":"/tmp/x"},
` + strings.Replace(aReplacingProcess, "%s", `,"replaces":["m.server"]`, 1) + `]}`
if _, err := Parse([]byte(raw)); err == nil || !strings.Contains(err.Error(), "still declares") {
t.Fatalf("a process replacing a declared resource was accepted: %v", err)
}
}
func TestOnlyAProcessThatStaysUpReplacesAnything(t *testing.T) {
for _, mode := range []string{`,"run-once":true`, `,"schedule":"0 3 * * *"`} {
if _, err := Parse([]byte(replacing(mode + `,"replaces":["m.server"]`))); err == nil {
t.Errorf("a process with %s was allowed to replace something", mode)
}
}
}
func TestAProcessCannotReplaceItselfOrNothing(t *testing.T) {
for _, bad := range []string{`,"replaces":["m.controller"]`, `,"replaces":[""]`} {
if _, err := Parse([]byte(replacing(bad))); err == nil {
t.Errorf("replaces %s was accepted", bad)
}
}
}
func TestOneThingHasOneReplacement(t *testing.T) {
second := strings.NewReplacer(`"id":"m.controller"`, `"id":"m.other"`, `"name":"m"`, `"name":"other"`).
Replace(strings.Replace(aReplacingProcess, "%s", `,"replaces":["m.server"]`, 1))
raw := `{"declaration":1,"resources":[` +
strings.Replace(aReplacingProcess, "%s", `,"replaces":["m.server"]`, 1) + "," + second + `]}`
if _, err := Parse([]byte(raw)); err == nil {
t.Fatal("two processes replacing one resource were accepted")
}
}
+49
View File
@@ -0,0 +1,49 @@
package declaration
import (
"strings"
"testing"
)
// novox/hq ADR 0177: a unit is in the system manager or an account's own; a user-scoped one names
// the account, a system one may not, and any other word is refused.
func TestAUserScopedUnitNamesItsAccountAndASystemOneMayNot(t *testing.T) {
cases := []struct{ scope, user, wants string }{
{"user", "ops", ""},
{"", "", ""},
{"system", "", ""},
{"user", "", "names the account"},
{"", "ops", "names no user"},
{"session", "ops", "scope \"session\""},
}
for _, c := range cases {
s := &Service{ID: "m.u", Type: TypeService, Unit: "u.service", State: "running", Scope: c.scope, User: c.user}
problems := s.validate("m.u", false)
got := strings.Join(problems, "; ")
if c.wants == "" && len(problems) != 0 {
t.Fatalf("scope %q user %q refused: %s", c.scope, c.user, got)
}
if c.wants != "" && !strings.Contains(got, c.wants) {
t.Fatalf("scope %q user %q: wanted a refusal saying %q, got %q", c.scope, c.user, c.wants, got)
}
}
}
// novox/hq ADR 0177: lingering is a field of the account, absent meaning nothing asserted.
func TestAnAccountMayBeDeclaredToLinger(t *testing.T) {
d, err := Parse([]byte(`{"declaration":1,"resources":[{"id":"m.login","type":"user","name":"ops","linger":true}]}`))
if err != nil {
t.Fatal(err)
}
u, ok := d.Resources[0].(*User)
if !ok || u.Linger == nil || !*u.Linger {
t.Fatalf("linger not read: %+v", d.Resources[0])
}
d, err = Parse([]byte(`{"declaration":1,"resources":[{"id":"m.login","type":"user","name":"ops"}]}`))
if err != nil {
t.Fatal(err)
}
if u := d.Resources[0].(*User); u.Linger != nil {
t.Fatal("an account that said nothing about lingering asserts it")
}
}
+326
View File
@@ -0,0 +1,326 @@
package firewall
import (
"context"
"fmt"
"regexp"
"sort"
"strings"
)
// What filters a machine, said with an owner (novox/hq ADR 0168).
//
// "The firewall found" names one front end, and a machine carries rules from several sources: the
// front end's own, the container runtime's plumbing, a ban list, the mesh's own tables, and whatever
// a predecessor installed directly — on both machines of the first mesh, in the user chain the
// runtime leaves for an administrator, where the mesh's reader of rules counted it as the runtime's.
// So the host reports every table and chain that refuses traffic, each with whose it is, and the
// mesh says truthfully what filters a converged machine. It removes none of it.
// Owners of a refusal.
const (
// OwnerMesh is the mesh's own tables: the derived filter and the guard.
OwnerMesh = "mesh"
// OwnerFoundFirewall is the front end found on the machine — ufw's chains.
OwnerFoundFirewall = "found-firewall"
// OwnerRuntime is the container runtime's own plumbing: its chains, the forward policy it sets
// when it turns forwarding on, its guard against reaching a container's address from off its
// bridge. Not the user chain it leaves for an administrator.
OwnerRuntime = "runtime"
// OwnerBan is a refusal that names the sources it refuses, in a chain that accepts nothing — a
// ban list, which is not a firewall.
OwnerBan = "ban"
// OwnerOther is everything else: rules the mesh did not write and cannot attribute. Where a
// predecessor's rules live.
OwnerOther = "other"
)
// A Filter is one place on the machine that refuses traffic: a chain of a table, or a chain of the
// legacy filter, with its owner and what it refuses in one line.
type Filter struct {
// Where names the chain: "table ip filter, chain DOCKER-USER", or "chain HAL-MESH-ONLY
// (iptables-legacy)".
Where string `json:"where"`
// Owner is one of the owners above.
Owner string `json:"owner"`
// Refuses is the first refusing line, counters stripped, and how many more there are.
Refuses string `json:"refuses"`
table, chain string
}
// userChain is the chain the container runtime creates empty and leaves for an administrator's
// rules, consulted before its own forwarding. Nothing in it is the runtime's.
const userChain = "DOCKER-USER"
// Filters classifies every refusing chain of an `nft list ruleset` and of the legacy filter's `-S`
// listings (by tool: iptables-legacy, ip6tables-legacy), in the order they appear.
func Filters(ruleset string, legacy map[string]string, ufwActive bool) []Filter {
var out []Filter
r := parseNft(ruleset)
refusing := map[string][]nftRule{} // by "table\x00chain"
for _, rule := range r.refusals {
k := rule.table + "\x00" + rule.chain
refusing[k] = append(refusing[k], rule)
}
for _, k := range r.chainOrder {
c := r.chains[k]
table, chain, _ := strings.Cut(k, "\x00")
rules := refusing[k]
if !c.dropping && len(rules) == 0 {
continue
}
f := Filter{table: table, chain: chain, Where: "table " + table + ", chain " + chain}
switch {
case table == MeshTable || table == "inet mesh_guard":
f.Owner = OwnerMesh
case strings.HasPrefix(chain, "ufw"):
f.Owner = OwnerFoundFirewall
if !ufwActive {
// Left behind by a retired front end, and still refusing: not ufw's any more in
// any sense that matters, since nothing maintains it.
f.Owner = OwnerOther
}
case chain == userChain:
f.Owner = OwnerOther
case c.dropping && (r.managed[table] || iptablesTable(table)) && runtimes(table, chain, c.policyLine):
f.Owner = OwnerRuntime
case len(rules) > 0 && (r.managed[table] || iptablesTable(table)) && allRuntimes(table, chain, rules):
f.Owner = OwnerRuntime
case len(rules) > 0 && allBans(r, rules):
f.Owner = OwnerBan
case c.dropping && !iptablesTable(table) && !r.managed[table] && len(rules) == 0:
// A table of its own whose base chain drops by policy: a firewall nobody declared.
f.Owner = OwnerOther
default:
f.Owner = OwnerOther
}
if ufwActive && (r.managed[table] || iptablesTable(table)) && f.Owner == OwnerOther && len(rules) == 0 && c.dropping {
// A base chain ufw set to drop while it is in force is ufw's.
f.Owner = OwnerFoundFirewall
}
f.Refuses = refusesLine(c, rules)
out = append(out, f)
}
tools := make([]string, 0, len(legacy))
for tool := range legacy {
tools = append(tools, tool)
}
sort.Strings(tools)
for _, tool := range tools {
out = append(out, legacyFilters(legacy[tool], tool, ufwActive)...)
}
return out
}
// allRuntimes is whether every refusal in a chain is the runtime's own.
func allRuntimes(table, chain string, rules []nftRule) bool {
for _, rule := range rules {
if !runtimes(table, chain, rule.line) {
return false
}
}
return true
}
// allBans is whether every refusal in a chain only bans the sources it names.
func allBans(r *nftRuleset, rules []nftRule) bool {
for _, rule := range rules {
if !r.onlyBans(rule) {
return false
}
}
return true
}
var counters = regexp.MustCompile(`\s*counter packets \d+ bytes \d+`)
// refusesLine is one line a person reads: the policy when the chain drops by policy, else the first
// refusing rule with its counters stripped, and how many more there are.
func refusesLine(c *nftChain, rules []nftRule) string {
var parts []string
if c.dropping {
parts = append(parts, "policy drop")
}
if len(rules) > 0 {
line := strings.TrimSpace(counters.ReplaceAllString(rules[0].line, ""))
if len(rules) > 1 {
line += fmt.Sprintf(" (and %d more)", len(rules)-1)
}
parts = append(parts, line)
}
return strings.Join(parts, "; ")
}
// legacyFilters classifies the chains of an `iptables-legacy -S` listing that refuse.
func legacyFilters(rules, tool string, ufwActive bool) []Filter {
policy := map[string]string{}
accepting := map[string]bool{}
jumpedFrom := map[string][]string{}
for _, line := range strings.Split(rules, "\n") {
fields := strings.Fields(line)
if len(fields) < 3 {
continue
}
switch fields[0] {
case "-P":
policy[fields[1]] = fields[2]
case "-A":
for i, f := range fields {
if (f == "-j" || f == "-g") && i+1 < len(fields) {
switch fields[i+1] {
case "ACCEPT":
accepting[fields[1]] = true
case "DROP", "REJECT", "RETURN", "LOG":
default:
jumpedFrom[fields[i+1]] = append(jumpedFrom[fields[i+1]], fields[1])
}
}
}
}
}
// A chain of refusals is a ban list when every refusal names the sources it refuses and the
// chain accepts nothing — the same rule the nftables side applies, and no more.
//
// **The policy of the chains that jump to it says nothing about what it is.** An earlier cut
// required every path into the chain to come from a built-in whose policy accepts, and the
// mesh's own intrusion prevention then read as a foreign rule set on the home server: its ban
// chain hangs off the container runtime's user chain as well as INPUT, and that machine's
// forward policy is DROP because the runtime set it. The machine reported "NOT the mesh alone"
// about a chain the mesh had just written (novox/hq ADR 0186). The policy is already classified
// where it belongs — as the runtime's — so requiring it here counted it twice.
ban := func(chain, line string) bool {
return bansSources(line) && !accepting[chain] && len(jumpedFrom[chain]) > 0
}
type seen struct {
owner string
lines []string
}
chains := map[string]*seen{}
var order []string
note := func(chain, owner, line string) {
s := chains[chain]
if s == nil {
s = &seen{owner: owner}
chains[chain] = s
order = append(order, chain)
}
if owner == OwnerOther || s.owner == "" {
s.owner = owner
}
s.lines = append(s.lines, line)
}
for _, line := range strings.Split(rules, "\n") {
fields := strings.Fields(line)
if len(fields) < 3 {
continue
}
chain := fields[1]
switch fields[0] {
case "-P":
if fields[2] != "DROP" {
continue
}
owner := OwnerOther
if chain == "FORWARD" {
owner = OwnerRuntime
}
if ufwActive {
owner = OwnerFoundFirewall
}
note(chain, owner, "policy DROP")
case "-A":
refuses := false
for i, f := range fields {
if f == "-j" && i+1 < len(fields) && (fields[i+1] == "DROP" || fields[i+1] == "REJECT") {
refuses = true
}
}
if !refuses {
continue
}
owner := OwnerOther
switch {
case strings.HasPrefix(chain, "ufw"):
owner = OwnerFoundFirewall
if !ufwActive {
owner = OwnerOther
}
case chain != userChain && strings.HasPrefix(chain, "DOCKER"):
owner = OwnerRuntime
case ban(chain, line):
owner = OwnerBan
}
note(chain, owner, strings.TrimSpace(line))
}
}
var out []Filter
for _, chain := range order {
s := chains[chain]
refuses := s.lines[0]
if len(s.lines) > 1 {
refuses += fmt.Sprintf(" (and %d more)", len(s.lines)-1)
}
out = append(out, Filter{Where: "chain " + chain + " (" + tool + ")", Owner: s.owner, Refuses: refuses})
}
return out
}
// Collect reads what filters this machine now: its nftables ruleset and, where the legacy tools
// exist, their listings. A machine without nft is read through iptables, as Detect reads it.
func Collect(ctx context.Context, run Runner, ufwActive bool) ([]Filter, error) {
ruleset := ""
noNft := false
out, err := run(ctx, "nft", "list", "ruleset")
switch {
case err == nil:
ruleset = out
case missing(err):
noNft = true
default:
return nil, fmt.Errorf("cannot read this machine's packet filter: %w", err)
}
legacy := map[string]string{}
tools := []string{"iptables-legacy", "ip6tables-legacy"}
if noNft {
tools = append(tools, "iptables", "ip6tables")
}
for _, tool := range tools {
if out, err := run(ctx, tool, "-S"); err == nil && strings.TrimSpace(out) != "" {
legacy[tool] = out
}
}
return Filters(ruleset, legacy, ufwActive), nil
}
// Alone is whether a machine is filtered by the mesh alone: nothing in the list but the mesh's
// own tables, the runtime's plumbing and bans (novox/hq ADR 0168).
func Alone(filters []Filter) bool {
for _, f := range filters {
if f.Owner == OwnerOther || f.Owner == OwnerFoundFirewall {
return false
}
}
return true
}
// Active says whether ufw is in force on this machine now. A machine without ufw is not.
func Active(ctx context.Context, run Runner) bool {
out, err := run(ctx, "ufw", "status")
return err == nil && statusActive(out)
}
// Installed says whether ufw is on this machine at all: a command that is not there is a front end
// that was uninstalled (novox/hq ADR 0180), not one that is silent.
func Installed(ctx context.Context, run Runner) bool {
_, err := run(ctx, "ufw", "status")
return !missing(err)
}
// Retirements of a found firewall, as the host records them.
const (
RetiredByMesh = "mesh"
RetiredFoundSo = "found-inactive"
// RetiredRemoved is a front end uninstalled by the module that replaced it (ADR 0180).
RetiredRemoved = "removed"
)
+130
View File
@@ -0,0 +1,130 @@
package firewall
import (
"os"
"strings"
"testing"
)
func fixture(t *testing.T, name string) string {
t.Helper()
raw, err := os.ReadFile("testdata/" + name)
if err != nil {
t.Fatal(err)
}
return string(raw)
}
func ownerOf(filters []Filter, where string) string {
for _, f := range filters {
if f.Where == where {
return f.Owner
}
}
return "(not reported)"
}
// Every refusing table and chain is classified with an owner (novox/hq ADR 0168), over rulesets
// captured from three machines of the first mesh. The control node: a ban list reached through the
// runtime's user chain is a ban; a refusal left in that chain, and a chain a retired front end left
// behind, are *other*; the runtime's own and the mesh's own are theirs.
func TestTheControlNodesRefusalsAreClassified(t *testing.T) {
got := Filters(fixture(t, "control-node.nft"), nil, false)
for where, want := range map[string]string{
"table ip filter, chain f2b-recidive": OwnerBan,
"table ip filter, chain DOCKER": OwnerRuntime,
"table ip raw, chain PREROUTING": OwnerRuntime,
"table inet mesh, chain input": OwnerMesh,
"table inet mesh, chain forward": OwnerMesh,
"table ip6 filter, chain DOCKER-USER": OwnerOther,
"table ip6 filter, chain ufw6-docker-logging-deny": OwnerOther,
} {
if o := ownerOf(got, where); o != want {
t.Errorf("%s: %s, want %s", where, o, want)
}
}
if Alone(got) {
t.Error("a machine with a refusal in the runtime's user chain reads as filtered by the mesh alone")
}
// What refuses adoption does not move (rule 4): the user chain's refusals are reported, not
// refused. The chain a retired front end left behind, still dropping, is what it always was
// to Detect — a refusal nobody speaks for, in one table.
if refusing := Refusing(fixture(t, "control-node.nft"), false); len(refusing) != 1 || refusing[0] != "table ip6 filter" {
t.Errorf("adoption's threshold moved: %v", refusing)
}
// The counters are stripped from what a person reads.
for _, f := range got {
if strings.Contains(f.Refuses, "counter packets") {
t.Errorf("counters in the line: %s", f.Refuses)
}
}
}
// The laptop: the runtime's forward policy and bridge guards, a virtualisation host and an endpoint
// agent that refuse nothing, and the mesh — filtered by the mesh alone.
func TestTheLaptopIsFilteredByTheMeshAlone(t *testing.T) {
got := Filters(fixture(t, "laptop.nft"), nil, false)
for where, want := range map[string]string{
"table ip filter, chain FORWARD": OwnerRuntime,
"table ip filter, chain DOCKER": OwnerRuntime,
"table ip raw, chain PREROUTING": OwnerRuntime,
"table inet mesh, chain input": OwnerMesh,
} {
if o := ownerOf(got, where); o != want {
t.Errorf("%s: %s, want %s", where, o, want)
}
}
for _, f := range got {
if strings.Contains(f.Where, "incus") || strings.Contains(f.Where, "fct_") {
t.Errorf("a table that refuses nothing is reported: %+v", f)
}
}
if !Alone(got) {
t.Errorf("the laptop is not read as filtered by the mesh alone: %+v", got)
}
}
// The home server: its rules are in the legacy filter, where a predecessor's chain still drops what
// arrives on the outward link for the forwarded path — invisible to the mesh until now (issue 144).
func TestThePredecessorsChainInTheLegacyFilterIsOther(t *testing.T) {
mesh := "table inet mesh {\n\tchain forward {\n\t\ttype filter hook forward priority filter; policy drop;\n\t}\n}\n"
got := Filters(mesh, map[string]string{"iptables-legacy": fixture(t, "home-server-legacy-S.txt")}, false)
for where, want := range map[string]string{
"table inet mesh, chain forward": OwnerMesh,
"chain FORWARD (iptables-legacy)": OwnerRuntime,
"chain DOCKER (iptables-legacy)": OwnerRuntime,
"chain HAL-MESH-ONLY (iptables-legacy)": OwnerOther,
} {
if o := ownerOf(got, where); o != want {
t.Errorf("%s: %s, want %s", where, o, want)
}
}
var other Filter
for _, f := range got {
if f.Owner == OwnerOther {
other = f
}
}
if !strings.Contains(other.Refuses, "-j DROP") {
t.Errorf("what the predecessor's chain refuses is not said: %+v", other)
}
if Alone(got) {
t.Error("a machine with a predecessor's chain reads as filtered by the mesh alone")
}
}
// With the front end in force, its chains are its own; retired, a chain it left behind that still
// refuses is nobody's and said so.
func TestAFrontEndsChainsAreItsWhileItIsInForce(t *testing.T) {
ruleset := dockerOnly(t) + ufwChains
for _, f := range Filters(ruleset, nil, true) {
if strings.Contains(f.Where, "ufw") && f.Owner != OwnerFoundFirewall {
t.Errorf("active: %+v", f)
}
}
for _, f := range Filters(ruleset, nil, false) {
if strings.Contains(f.Where, "ufw") && f.Owner != OwnerOther {
t.Errorf("retired: %+v", f)
}
}
}
+90 -69
View File
@@ -128,42 +128,82 @@ func statusActive(out string) bool {
// mesh needs, so a refusal that names the sources it refuses, in a table or a chain that accepts // mesh needs, so a refusal that names the sources it refuses, in a table or a chain that accepts
// nothing and is entered only from chains whose policy accepts, is not counted. // nothing and is entered only from chains whose policy accepts, is not counted.
func Refusing(ruleset string, ufwActive bool) []string { func Refusing(ruleset string, ufwActive bool) []string {
type rule struct{ table, chain, line string } var refusing []string
type chainOf struct { for _, f := range Filters(ruleset, nil, ufwActive) {
if f.Owner != OwnerOther || f.chain == userChain {
// A refusal in the runtime's user chain is reported as *other* and does not refuse
// adoption (novox/hq ADR 0168, rule 4): both predecessors kept their rules there.
continue
}
name := "table " + f.table
if len(refusing) == 0 || refusing[len(refusing)-1] != name {
if !contains(refusing, name) {
refusing = append(refusing, name)
}
}
}
return refusing
}
func contains(list []string, s string) bool {
for _, x := range list {
if x == s {
return true
}
}
return false
}
// nftRule is one line of a ruleset that refuses, with where it is.
type nftRule struct{ table, chain, line string }
// nftChain is what a parse knows about one chain.
type nftChain struct {
base, dropping, accepts bool base, dropping, accepts bool
policyLine string policyLine string
jumpedFrom []string jumpedFrom []string
} }
chains := map[string]*chainOf{} // by "table\x00chain"
tableAccepts := map[string]bool{} // nftRuleset is `nft list ruleset`, read: its tables in order, its chains, every refusing line,
var tables []string // and which tables iptables-nft manages.
var refusals []rule type nftRuleset struct {
managed := map[string]bool{} tables []string
var table, chain string chains map[string]*nftChain // by "table\x00chain"
get := func(t, c string) *chainOf { chainOrder []string
tableAccepts map[string]bool
refusals []nftRule
managed map[string]bool
}
func (r *nftRuleset) get(t, c string) *nftChain {
k := t + "\x00" + c k := t + "\x00" + c
if chains[k] == nil { if r.chains[k] == nil {
chains[k] = &chainOf{} r.chains[k] = &nftChain{}
r.chainOrder = append(r.chainOrder, k)
} }
return chains[k] return r.chains[k]
} }
func parseNft(ruleset string) *nftRuleset {
r := &nftRuleset{chains: map[string]*nftChain{}, tableAccepts: map[string]bool{}, managed: map[string]bool{}}
var table, chain string
for _, raw := range strings.Split(ruleset, "\n") { for _, raw := range strings.Split(ruleset, "\n") {
line := strings.TrimSpace(raw) line := strings.TrimSpace(raw)
switch { switch {
case strings.HasPrefix(line, "# Warning: table ") && strings.Contains(line, "managed by iptables-nft"): case strings.HasPrefix(line, "# Warning: table ") && strings.Contains(line, "managed by iptables-nft"):
name := strings.TrimPrefix(line, "# Warning: table ") name := strings.TrimPrefix(line, "# Warning: table ")
name, _, _ = strings.Cut(name, " is managed") name, _, _ = strings.Cut(name, " is managed")
managed[name] = true r.managed[name] = true
continue continue
case strings.HasPrefix(line, "table "): case strings.HasPrefix(line, "table "):
table = strings.TrimSuffix(strings.TrimSpace(strings.TrimPrefix(line, "table ")), "{") table = strings.TrimSuffix(strings.TrimSpace(strings.TrimPrefix(line, "table ")), "{")
table = strings.TrimSpace(table) table = strings.TrimSpace(table)
tables = append(tables, table) r.tables = append(r.tables, table)
chain = "" chain = ""
continue continue
case strings.HasPrefix(line, "chain "): case strings.HasPrefix(line, "chain "):
chain = strings.TrimSpace(strings.TrimSuffix(strings.TrimPrefix(line, "chain "), "{")) chain = strings.TrimSpace(strings.TrimSuffix(strings.TrimPrefix(line, "chain "), "{"))
get(table, chain) r.get(table, chain)
continue continue
case strings.HasPrefix(line, "set ") || strings.HasPrefix(line, "map ") || case strings.HasPrefix(line, "set ") || strings.HasPrefix(line, "map ") ||
strings.HasPrefix(line, "flowtable "): strings.HasPrefix(line, "flowtable "):
@@ -172,7 +212,7 @@ func Refusing(ruleset string, ufwActive bool) []string {
case line == "" || line == "}" || strings.HasPrefix(line, "#") || chain == "": case line == "" || line == "}" || strings.HasPrefix(line, "#") || chain == "":
continue continue
} }
c := get(table, chain) c := r.get(table, chain)
if strings.HasPrefix(line, "type ") { if strings.HasPrefix(line, "type ") {
c.base = true c.base = true
c.policyLine = line c.policyLine = line
@@ -183,87 +223,68 @@ func Refusing(ruleset string, ufwActive bool) []string {
if i := strings.Index(line, verb); i >= 0 { if i := strings.Index(line, verb); i >= 0 {
target := strings.Fields(line[i+len(verb):]) target := strings.Fields(line[i+len(verb):])
if len(target) > 0 { if len(target) > 0 {
get(table, target[0]).jumpedFrom = append(get(table, target[0]).jumpedFrom, chain) r.get(table, target[0]).jumpedFrom = append(r.get(table, target[0]).jumpedFrom, chain)
} }
} }
} }
if accepts(line) { if accepts(line) {
c.accepts = true c.accepts = true
tableAccepts[table] = true r.tableAccepts[table] = true
} }
if verdictRefuses(line) { if verdictRefuses(line) {
refusals = append(refusals, rule{table, chain, line}) r.refusals = append(r.refusals, nftRule{table, chain, line})
} }
} }
return r
}
skipped := func(table string) bool { // onlyBans is whether a refusal only refuses the sources it names: in a table that accepts nothing
if table == "inet mesh" || table == "inet mesh_guard" { // and whose base chains all accept by default, or in a chain that accepts nothing and is entered
return true // only from base chains that accept by default.
} func (r *nftRuleset) onlyBans(rule nftRule) bool {
return (managed[table] || iptablesTable(table)) && ufwActive if !bansSources(rule.line) {
}
// onlyBans is whether a refusal only refuses the sources it names: in a table that accepts
// nothing and whose base chains all accept by default, or in a chain that accepts nothing and
// is entered only from base chains that accept by default.
onlyBans := func(r rule) bool {
if !bansSources(r.line) {
return false return false
} }
allAccepting := true allAccepting := true
for k, c := range chains { for k, c := range r.chains {
if strings.HasPrefix(k, r.table+"\x00") && c.base && !strings.Contains(c.policyLine, "policy accept") { if strings.HasPrefix(k, rule.table+"\x00") && c.base && !strings.Contains(c.policyLine, "policy accept") {
allAccepting = false allAccepting = false
} }
} }
if !tableAccepts[r.table] && allAccepting { if !r.tableAccepts[rule.table] && allAccepting {
return true return true
} }
c := get(r.table, r.chain) return r.enteredAccepting(rule.table, rule.chain, map[string]bool{})
}
// enteredAccepting is whether a chain accepts nothing and is entered only through chains that
// accept by default — base chains whose policy accepts, or chains that are themselves entered that
// way and accept nothing. A ban list jumped to from the runtime's user chain, which the forward
// chain enters with an accepting policy, is still a ban list.
func (r *nftRuleset) enteredAccepting(table, chain string, seen map[string]bool) bool {
if seen[chain] {
return false
}
seen[chain] = true
c := r.get(table, chain)
if c.base || c.accepts || len(c.jumpedFrom) == 0 { if c.base || c.accepts || len(c.jumpedFrom) == 0 {
return false return false
} }
for _, from := range c.jumpedFrom { for _, from := range c.jumpedFrom {
caller := get(r.table, from) caller := r.get(table, from)
if !caller.base || !strings.Contains(caller.policyLine, "policy accept") { if caller.base {
if !strings.Contains(caller.policyLine, "policy accept") {
return false
}
continue
}
if caller.accepts || !r.enteredAccepting(table, from, seen) {
return false return false
} }
} }
return true return true
} }
counted := map[string]bool{}
for k, c := range chains {
t, name, _ := strings.Cut(k, "\x00")
if skipped(t) || !c.dropping {
continue
}
if (managed[t] || iptablesTable(t)) && runtimes(t, name, c.policyLine) {
continue
}
counted[t] = true
}
for _, r := range refusals {
if skipped(r.table) || counted[r.table] {
continue
}
if (managed[r.table] || iptablesTable(r.table)) && runtimes(r.table, r.chain, r.line) {
continue
}
if onlyBans(r) {
continue
}
counted[r.table] = true
}
var refusing []string
for _, t := range tables {
if counted[t] {
counted[t] = false
refusing = append(refusing, "table "+t)
}
}
return refusing
}
// iptablesTable is whether a table is one iptables-nft writes. Named rather than read from the // iptablesTable is whether a table is one iptables-nft writes. Named rather than read from the
// warning nft prints above it, because nft does not print that for every such table: a captured // warning nft prints above it, because nft does not print that for every such table: a captured
// ruleset carried it on ip filter and not on ip raw, where the runtime keeps its drops. // ruleset carried it on ip filter and not on ip raw, where the runtime keeps its drops.
+60
View File
@@ -0,0 +1,60 @@
package firewall
import (
"os"
"testing"
)
// The mesh's own ban list is a ban wherever it hangs (novox/hq ADR 0186).
//
// Captured from the home server after the intrusion prevention had banned four addresses: its ban
// chain is jumped to from INPUT, whose policy accepts, and from the container runtime's user chain,
// which hangs off a FORWARD the runtime set to DROP. Requiring every path to come from an accepting
// built-in made the machine report "NOT the mesh alone" about a chain the mesh had just written.
func TestTheMeshsOwnBanChainIsABanBehindADroppingForward(t *testing.T) {
legacy, err := os.ReadFile("testdata/home-server-bans-S.txt")
if err != nil {
t.Fatal(err)
}
filters := Filters("", map[string]string{"iptables-legacy": string(legacy)}, false)
var ban, other []string
for _, f := range filters {
switch f.Owner {
case OwnerBan:
ban = append(ban, f.Where)
case OwnerOther:
other = append(other, f.Where)
}
}
if len(other) > 0 {
t.Errorf("the machine reports %v as rule sets the mesh did not write", other)
}
found := false
for _, w := range ban {
if w == "chain f2b-route-proxy (iptables-legacy)" {
found = true
}
}
if !found {
t.Errorf("the intrusion prevention's own chain was not read as a ban; bans were %v", ban)
}
if !Alone(filters) {
t.Error("a machine filtered by the mesh and its own bans does not read as the mesh alone")
}
}
// A chain that accepts anything is doing more than banning, and is still not a ban — which is what
// keeps a predecessor's allow-and-drop chain classified as something the operator must look at.
func TestAChainThatAcceptsIsNotABan(t *testing.T) {
rules := "-P INPUT ACCEPT\n" +
"-A INPUT -j HAL-MESH-ONLY\n" +
"-N HAL-MESH-ONLY\n" +
"-A HAL-MESH-ONLY -s 10.0.0.0/8 -j ACCEPT\n" +
"-A HAL-MESH-ONLY -s 203.0.113.7/32 -j DROP\n"
for _, f := range Filters("", map[string]string{"iptables-legacy": rules}, false) {
if f.Where == "chain HAL-MESH-ONLY (iptables-legacy)" && f.Owner != OwnerOther {
t.Errorf("a chain that accepts was classified as %s", f.Owner)
}
}
}
+588
View File
@@ -0,0 +1,588 @@
# Warning: table ip filter is managed by iptables-nft, do not touch!
table ip filter {
chain INPUT {
type filter hook input priority filter; policy accept;
ip protocol tcp counter packets 945757787 bytes 1737008792038 jump f2b-sshd
ip protocol tcp counter packets 945756610 bytes 1737008898620 jump f2b-recidive
counter packets 2862213204 bytes 3144751431654 jump ufw-before-logging-input
counter packets 2862213204 bytes 3144751431654 jump ufw-before-input
counter packets 989333889 bytes 1776344988272 jump ufw-after-input
counter packets 989303248 bytes 1776343408920 jump ufw-after-logging-input
counter packets 989303248 bytes 1776343408920 jump ufw-reject-input
counter packets 989303248 bytes 1776343408920 jump ufw-track-input
}
chain FORWARD {
type filter hook forward priority filter; policy accept;
oifname "mesh0" counter packets 1613103 bytes 2577614868 accept
iifname "mesh0" counter packets 995195 bytes 84526284 accept
counter packets 20454697 bytes 11504107676 jump DOCKER-USER
counter packets 20442192 bytes 11503368404 jump DOCKER-FORWARD
counter packets 12438285 bytes 10907281833 jump ufw-before-logging-forward
counter packets 12438285 bytes 10907281833 jump ufw-before-forward
counter packets 384 bytes 39643 jump ufw-after-forward
counter packets 384 bytes 39643 jump ufw-after-logging-forward
counter packets 384 bytes 39643 jump ufw-reject-forward
counter packets 384 bytes 39643 jump ufw-track-forward
}
chain OUTPUT {
type filter hook output priority filter; policy accept;
counter packets 3195070897 bytes 3951725261199 jump ufw-before-logging-output
counter packets 3195070897 bytes 3951725261199 jump ufw-before-output
counter packets 945745931 bytes 1778546547406 jump ufw-after-output
counter packets 945745931 bytes 1778546547406 jump ufw-after-logging-output
counter packets 945745931 bytes 1778546547406 jump ufw-reject-output
counter packets 945745931 bytes 1778546547406 jump ufw-track-output
}
chain DOCKER-FORWARD {
counter packets 20442192 bytes 11503368404 jump DOCKER-CT
counter packets 8079126 bytes 1230017985 jump DOCKER-INTERNAL
counter packets 8079126 bytes 1230017985 jump DOCKER-BRIDGE
iifname "br-cadedce55fe9" counter packets 0 bytes 0 accept
iifname "br-dd007c7e67bc" counter packets 0 bytes 0 accept
iifname "br-a5fbc29c2c2a" counter packets 0 bytes 0 accept
iifname "br-6eb1e7f7f847" counter packets 0 bytes 0 accept
iifname "br-8ce143481a5b" counter packets 14700 bytes 2493600 accept
iifname "br-84e7d0cfeada" counter packets 0 bytes 0 accept
iifname "br-f8b083119d99" counter packets 264 bytes 57438 accept
iifname "br-0d1490cc67c9" counter packets 732468 bytes 351624109 accept
iifname "br-3b338a381229" counter packets 137 bytes 11876 accept
iifname "br-3008d408e73a" counter packets 25380 bytes 1564417 accept
iifname "br-ca07a9577a7f" counter packets 0 bytes 0 accept
iifname "docker0" counter packets 6695791 bytes 795364128 accept
iifname "br-a63fa64a9e18" counter packets 0 bytes 0 accept
iifname "br-1ccb887b3344" counter packets 237174 bytes 36804979 accept
iifname "br-9fd22324ec08" counter packets 0 bytes 0 accept
iifname "br-73641cceafc3" counter packets 36 bytes 6614 accept
iifname "br-77eb8a9e2ba1" counter packets 0 bytes 0 accept
iifname "br-e99ce5248c84" counter packets 0 bytes 0 accept
iifname "br-e5d78502832d" counter packets 0 bytes 0 accept
iifname "br-2e4a76a7cd2e" counter packets 20339 bytes 1799711 accept
iifname "br-72fd626a8ff7" counter packets 0 bytes 0 accept
}
chain DOCKER-USER {
ip protocol tcp counter packets 3054221 bytes 3329963005 jump f2b-sshd
ip protocol tcp counter packets 3054221 bytes 3329963005 jump f2b-recidive
counter packets 1421378091 bytes 2045004412819 return
}
chain ufw-before-logging-input {
}
chain ufw-before-logging-output {
}
chain ufw-before-logging-forward {
}
chain ufw-before-input {
}
chain ufw-before-output {
}
chain ufw-before-forward {
}
chain ufw-after-input {
}
chain ufw-after-output {
}
chain ufw-after-forward {
}
chain ufw-after-logging-input {
}
chain ufw-after-logging-output {
}
chain ufw-after-logging-forward {
}
chain ufw-reject-input {
}
chain ufw-reject-output {
}
chain ufw-reject-forward {
}
chain ufw-track-input {
}
chain ufw-track-output {
}
chain ufw-track-forward {
}
chain DOCKER {
ip daddr 172.17.0.7 iifname != "docker0" oifname "docker0" tcp dport 5432 counter packets 8 bytes 480 accept
ip daddr 172.19.0.2 iifname != "br-72fd626a8ff7" oifname "br-72fd626a8ff7" tcp dport 8080 counter packets 0 bytes 0 accept
ip daddr 172.17.0.6 iifname != "docker0" oifname "docker0" tcp dport 9443 counter packets 0 bytes 0 accept
ip daddr 172.17.0.6 iifname != "docker0" oifname "docker0" tcp dport 9000 counter packets 0 bytes 0 accept
ip daddr 192.168.176.2 iifname != "br-f8b083119d99" oifname "br-f8b083119d99" tcp dport 9001 counter packets 0 bytes 0 accept
ip daddr 192.168.176.2 iifname != "br-f8b083119d99" oifname "br-f8b083119d99" tcp dport 9000 counter packets 47769 bytes 2866140 accept
ip daddr 172.20.0.2 iifname != "br-6eb1e7f7f847" oifname "br-6eb1e7f7f847" tcp dport 8080 counter packets 0 bytes 0 accept
ip daddr 172.27.0.2 iifname != "br-3008d408e73a" oifname "br-3008d408e73a" tcp dport 3000 counter packets 0 bytes 0 accept
ip daddr 192.168.48.5 iifname != "br-77eb8a9e2ba1" oifname "br-77eb8a9e2ba1" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 192.168.48.4 iifname != "br-77eb8a9e2ba1" oifname "br-77eb8a9e2ba1" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 192.168.48.3 iifname != "br-77eb8a9e2ba1" oifname "br-77eb8a9e2ba1" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 192.168.48.2 iifname != "br-77eb8a9e2ba1" oifname "br-77eb8a9e2ba1" tcp dport 9000 counter packets 0 bytes 0 accept
ip daddr 172.18.0.2 iifname != "br-2e4a76a7cd2e" oifname "br-2e4a76a7cd2e" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 172.17.0.5 iifname != "docker0" oifname "docker0" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 172.17.0.3 iifname != "docker0" oifname "docker0" tcp dport 8222 counter packets 0 bytes 0 accept
ip daddr 172.17.0.3 iifname != "docker0" oifname "docker0" tcp dport 4222 counter packets 97 bytes 5744 accept
ip daddr 172.28.0.2 iifname != "br-8ce143481a5b" oifname "br-8ce143481a5b" tcp dport 1433 counter packets 0 bytes 0 accept
ip daddr 192.168.80.2 iifname != "br-e99ce5248c84" oifname "br-e99ce5248c84" tcp dport 8080 counter packets 0 bytes 0 accept
ip daddr 192.168.112.3 iifname != "br-e5d78502832d" oifname "br-e5d78502832d" tcp dport 9000 counter packets 0 bytes 0 accept
ip daddr 192.168.112.2 iifname != "br-e5d78502832d" oifname "br-e5d78502832d" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 192.168.128.2 iifname != "br-73641cceafc3" oifname "br-73641cceafc3" tcp dport 27017 counter packets 14 bytes 840 accept
ip daddr 192.168.208.2 iifname != "br-9fd22324ec08" oifname "br-9fd22324ec08" tcp dport 35621 counter packets 0 bytes 0 accept
ip daddr 192.168.203.13 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 4243 counter packets 0 bytes 0 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 995 counter packets 194 bytes 11000 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 993 counter packets 188 bytes 9394 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 587 counter packets 444 bytes 23312 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 465 counter packets 104 bytes 5852 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 443 counter packets 0 bytes 0 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 143 counter packets 443 bytes 25280 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 110 counter packets 192 bytes 9561 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 80 counter packets 0 bytes 0 accept
ip daddr 192.168.203.12 iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" tcp dport 25 counter packets 430 bytes 22919 accept
ip daddr 172.17.0.2 iifname != "docker0" oifname "docker0" tcp dport 3000 counter packets 0 bytes 0 accept
ip daddr 172.17.0.2 iifname != "docker0" oifname "docker0" tcp dport 22 counter packets 1578 bytes 93884 accept
ip daddr 172.17.0.4 iifname != "docker0" oifname "docker0" tcp dport 5000 counter packets 25 bytes 1492 accept
iifname != "br-cadedce55fe9" oifname "br-cadedce55fe9" counter packets 0 bytes 0 drop
iifname != "br-dd007c7e67bc" oifname "br-dd007c7e67bc" counter packets 0 bytes 0 drop
iifname != "br-a5fbc29c2c2a" oifname "br-a5fbc29c2c2a" counter packets 0 bytes 0 drop
iifname != "br-6eb1e7f7f847" oifname "br-6eb1e7f7f847" counter packets 0 bytes 0 drop
iifname != "br-8ce143481a5b" oifname "br-8ce143481a5b" counter packets 0 bytes 0 drop
iifname != "br-84e7d0cfeada" oifname "br-84e7d0cfeada" counter packets 0 bytes 0 drop
iifname != "br-f8b083119d99" oifname "br-f8b083119d99" counter packets 0 bytes 0 drop
iifname != "br-0d1490cc67c9" oifname "br-0d1490cc67c9" counter packets 0 bytes 0 drop
iifname != "br-3b338a381229" oifname "br-3b338a381229" counter packets 0 bytes 0 drop
iifname != "br-3008d408e73a" oifname "br-3008d408e73a" counter packets 0 bytes 0 drop
iifname != "br-ca07a9577a7f" oifname "br-ca07a9577a7f" counter packets 0 bytes 0 drop
iifname != "docker0" oifname "docker0" counter packets 0 bytes 0 drop
iifname != "br-a63fa64a9e18" oifname "br-a63fa64a9e18" counter packets 0 bytes 0 drop
iifname != "br-1ccb887b3344" oifname "br-1ccb887b3344" counter packets 0 bytes 0 drop
iifname != "br-9fd22324ec08" oifname "br-9fd22324ec08" counter packets 0 bytes 0 drop
iifname != "br-73641cceafc3" oifname "br-73641cceafc3" counter packets 0 bytes 0 drop
iifname != "br-77eb8a9e2ba1" oifname "br-77eb8a9e2ba1" counter packets 0 bytes 0 drop
iifname != "br-e99ce5248c84" oifname "br-e99ce5248c84" counter packets 0 bytes 0 drop
iifname != "br-e5d78502832d" oifname "br-e5d78502832d" counter packets 0 bytes 0 drop
iifname != "br-2e4a76a7cd2e" oifname "br-2e4a76a7cd2e" counter packets 0 bytes 0 drop
iifname != "br-72fd626a8ff7" oifname "br-72fd626a8ff7" counter packets 0 bytes 0 drop
}
chain DOCKER-BRIDGE {
oifname "br-cadedce55fe9" counter packets 0 bytes 0 jump DOCKER
oifname "br-dd007c7e67bc" counter packets 0 bytes 0 jump DOCKER
oifname "br-a5fbc29c2c2a" counter packets 0 bytes 0 jump DOCKER
oifname "br-6eb1e7f7f847" counter packets 799 bytes 47940 jump DOCKER
oifname "br-8ce143481a5b" counter packets 0 bytes 0 jump DOCKER
oifname "br-84e7d0cfeada" counter packets 0 bytes 0 jump DOCKER
oifname "br-f8b083119d99" counter packets 98911 bytes 5934660 jump DOCKER
oifname "br-0d1490cc67c9" counter packets 69740 bytes 4118476 jump DOCKER
oifname "br-3b338a381229" counter packets 32 bytes 1920 jump DOCKER
oifname "br-3008d408e73a" counter packets 458 bytes 27480 jump DOCKER
oifname "br-ca07a9577a7f" counter packets 0 bytes 0 jump DOCKER
oifname "docker0" counter packets 87073 bytes 5223529 jump DOCKER
oifname "br-a63fa64a9e18" counter packets 1353 bytes 81180 jump DOCKER
oifname "br-1ccb887b3344" counter packets 7662 bytes 419862 jump DOCKER
oifname "br-9fd22324ec08" counter packets 173 bytes 10380 jump DOCKER
oifname "br-73641cceafc3" counter packets 162 bytes 9720 jump DOCKER
oifname "br-77eb8a9e2ba1" counter packets 94 bytes 5640 jump DOCKER
oifname "br-e99ce5248c84" counter packets 8 bytes 480 jump DOCKER
oifname "br-e5d78502832d" counter packets 26 bytes 1560 jump DOCKER
oifname "br-2e4a76a7cd2e" counter packets 7 bytes 420 jump DOCKER
oifname "br-72fd626a8ff7" counter packets 0 bytes 0 jump DOCKER
}
chain DOCKER-CT {
oifname "br-cadedce55fe9" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-dd007c7e67bc" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-a5fbc29c2c2a" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-6eb1e7f7f847" xt match "conntrack" counter packets 38458 bytes 6234236 accept
oifname "br-8ce143481a5b" xt match "conntrack" counter packets 60403 bytes 20478794 accept
oifname "br-84e7d0cfeada" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-f8b083119d99" xt match "conntrack" counter packets 1008024 bytes 206364436 accept
oifname "br-0d1490cc67c9" xt match "conntrack" counter packets 871134 bytes 1416174426 accept
oifname "br-3b338a381229" xt match "conntrack" counter packets 4649 bytes 2311375 accept
oifname "br-3008d408e73a" xt match "conntrack" counter packets 13415 bytes 1974731 accept
oifname "br-ca07a9577a7f" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "docker0" xt match "conntrack" counter packets 8909862 bytes 7892163419 accept
oifname "br-a63fa64a9e18" xt match "conntrack" counter packets 16688 bytes 6822829 accept
oifname "br-1ccb887b3344" xt match "conntrack" counter packets 461453 bytes 141913887 accept
oifname "br-9fd22324ec08" xt match "conntrack" counter packets 1677 bytes 427538 accept
oifname "br-73641cceafc3" xt match "conntrack" counter packets 417619 bytes 35343624 accept
oifname "br-77eb8a9e2ba1" xt match "conntrack" counter packets 125437 bytes 94155193 accept
oifname "br-e99ce5248c84" xt match "conntrack" counter packets 91 bytes 19173 accept
oifname "br-e5d78502832d" xt match "conntrack" counter packets 128653 bytes 40539569 accept
oifname "br-2e4a76a7cd2e" xt match "conntrack" counter packets 20257 bytes 158631592 accept
oifname "br-72fd626a8ff7" xt match "conntrack" counter packets 0 bytes 0 accept
}
chain DOCKER-INTERNAL {
}
chain f2b-recidive {
ip saddr 2.57.122.209 counter packets 0 bytes 0 xt target "REJECT"
ip saddr 2.57.122.76 counter packets 127 bytes 7600 xt target "REJECT"
ip saddr 195.178.110.228 counter packets 17 bytes 1000 xt target "REJECT"
ip saddr 2.57.122.74 counter packets 11 bytes 620 xt target "REJECT"
ip saddr 195.178.110.26 counter packets 56 bytes 3360 xt target "REJECT"
ip saddr 92.118.39.77 counter packets 2 bytes 80 xt target "REJECT"
ip saddr 92.118.39.71 counter packets 1 bytes 40 xt target "REJECT"
ip saddr 45.148.10.240 counter packets 0 bytes 0 xt target "REJECT"
ip saddr 195.178.110.30 counter packets 8 bytes 320 xt target "REJECT"
counter packets 948810608 bytes 1740338848565 return
}
chain f2b-sshd {
counter packets 948810709 bytes 1740338655735 return
}
}
# Warning: table ip6 filter is managed by iptables-nft, do not touch!
table ip6 filter {
chain INPUT {
type filter hook input priority filter; policy accept;
counter packets 5426360 bytes 34419159588 jump ufw6-before-logging-input
counter packets 5426360 bytes 34419159588 jump ufw6-before-input
counter packets 367982 bytes 3415126642 jump ufw6-after-input
counter packets 367982 bytes 3415126642 jump ufw6-after-logging-input
counter packets 367982 bytes 3415126642 jump ufw6-reject-input
counter packets 367982 bytes 3415126642 jump ufw6-track-input
}
chain FORWARD {
type filter hook forward priority filter; policy accept;
counter packets 0 bytes 0 jump DOCKER-USER
counter packets 0 bytes 0 jump DOCKER-FORWARD
counter packets 0 bytes 0 jump ufw6-before-logging-forward
counter packets 0 bytes 0 jump ufw6-before-forward
counter packets 0 bytes 0 jump ufw6-after-forward
counter packets 0 bytes 0 jump ufw6-after-logging-forward
counter packets 0 bytes 0 jump ufw6-reject-forward
counter packets 0 bytes 0 jump ufw6-track-forward
}
chain OUTPUT {
type filter hook output priority filter; policy accept;
counter packets 6004354 bytes 1866587952 jump ufw6-before-logging-output
counter packets 6004354 bytes 1866587952 jump ufw6-before-output
counter packets 2241898 bytes 639173314 jump ufw6-after-output
counter packets 2241898 bytes 639173314 jump ufw6-after-logging-output
counter packets 2241898 bytes 639173314 jump ufw6-reject-output
counter packets 2241898 bytes 639173314 jump ufw6-track-output
}
chain DOCKER-FORWARD {
counter packets 0 bytes 0 jump DOCKER-CT
counter packets 0 bytes 0 jump DOCKER-INTERNAL
counter packets 0 bytes 0 jump DOCKER-BRIDGE
}
chain DOCKER-USER {
counter packets 0 bytes 0 jump ufw6-user-forward
xt match "conntrack" counter packets 0 bytes 0 return
xt match "conntrack" counter packets 0 bytes 0 drop
iifname "docker0" oifname "docker0" counter packets 0 bytes 0 accept
ip6 saddr fd00::/8 counter packets 0 bytes 0 return
ip6 daddr fd00::/8 xt match "conntrack" counter packets 0 bytes 0 jump ufw6-docker-logging-deny
counter packets 0 bytes 0 return
}
chain ufw6-before-logging-input {
}
chain ufw6-before-logging-output {
}
chain ufw6-before-logging-forward {
}
chain ufw6-before-input {
}
chain ufw6-before-output {
}
chain ufw6-before-forward {
}
chain ufw6-after-input {
}
chain ufw6-after-output {
}
chain ufw6-after-forward {
}
chain ufw6-after-logging-input {
}
chain ufw6-after-logging-output {
}
chain ufw6-after-logging-forward {
}
chain ufw6-reject-input {
}
chain ufw6-reject-output {
}
chain ufw6-reject-forward {
}
chain ufw6-track-input {
}
chain ufw6-track-output {
}
chain ufw6-track-forward {
}
chain ufw6-user-forward {
}
chain ufw6-docker-logging-deny {
limit rate 3/minute burst 10 packets counter packets 0 bytes 0 xt target "LOG"
counter packets 0 bytes 0 drop
}
chain DOCKER {
}
chain DOCKER-BRIDGE {
}
chain DOCKER-CT {
}
chain DOCKER-INTERNAL {
}
}
# Warning: table ip nat is managed by iptables-nft, do not touch!
table ip nat {
chain PREROUTING {
type nat hook prerouting priority dstnat; policy accept;
xt match "addrtype" counter packets 10757093 bytes 647829087 jump DOCKER
}
chain OUTPUT {
type nat hook output priority dstnat; policy accept;
ip daddr != 127.0.0.0/8 xt match "addrtype" counter packets 128611 bytes 7705178 jump DOCKER
}
chain POSTROUTING {
type nat hook postrouting priority srcnat; policy accept;
ip saddr 172.19.0.0/16 oifname != "br-72fd626a8ff7" counter packets 0 bytes 0 xt target "MASQUERADE"
ip saddr 172.18.0.0/16 oifname != "br-2e4a76a7cd2e" counter packets 825 bytes 49500 xt target "MASQUERADE"
ip saddr 192.168.112.0/20 oifname != "br-e5d78502832d" counter packets 126 bytes 7560 xt target "MASQUERADE"
ip saddr 192.168.80.0/20 oifname != "br-e99ce5248c84" counter packets 0 bytes 0 xt target "MASQUERADE"
ip saddr 192.168.48.0/20 oifname != "br-77eb8a9e2ba1" counter packets 99 bytes 5940 xt target "MASQUERADE"
ip saddr 192.168.128.0/20 oifname != "br-73641cceafc3" counter packets 536 bytes 32160 xt target "MASQUERADE"
ip saddr 192.168.208.0/20 oifname != "br-9fd22324ec08" counter packets 2 bytes 120 xt target "MASQUERADE"
ip saddr 192.168.203.0/24 oifname != "br-1ccb887b3344" counter packets 37248 bytes 2854476 xt target "MASQUERADE"
ip saddr 192.168.64.0/20 oifname != "br-a63fa64a9e18" counter packets 353 bytes 21180 xt target "MASQUERADE"
ip saddr 172.17.0.0/16 oifname != "docker0" counter packets 99933 bytes 6001436 xt target "MASQUERADE"
ip saddr 172.21.0.0/16 oifname != "br-84e7d0cfeada" counter packets 2 bytes 128 xt target "MASQUERADE"
ip saddr 192.168.176.0/20 oifname != "br-f8b083119d99" counter packets 209 bytes 12644 xt target "MASQUERADE"
ip saddr 172.20.0.0/16 oifname != "br-6eb1e7f7f847" counter packets 699 bytes 42516 xt target "MASQUERADE"
ip saddr 172.28.0.0/16 oifname != "br-8ce143481a5b" counter packets 1934 bytes 116040 xt target "MASQUERADE"
ip saddr 172.27.0.0/16 oifname != "br-3008d408e73a" counter packets 2904 bytes 174240 xt target "MASQUERADE"
ip saddr 172.25.0.0/16 oifname != "br-cadedce55fe9" counter packets 0 bytes 0 xt target "MASQUERADE"
ip saddr 172.24.0.0/16 oifname != "br-3b338a381229" counter packets 385 bytes 23164 xt target "MASQUERADE"
ip saddr 192.168.224.0/20 oifname != "br-ca07a9577a7f" counter packets 0 bytes 0 xt target "MASQUERADE"
ip saddr 192.168.0.0/20 oifname != "br-a5fbc29c2c2a" counter packets 10 bytes 600 xt target "MASQUERADE"
ip saddr 172.31.0.0/16 oifname != "br-dd007c7e67bc" counter packets 0 bytes 0 xt target "MASQUERADE"
ip saddr 192.168.240.0/20 oifname != "br-0d1490cc67c9" counter packets 587045 bytes 35224163 xt target "MASQUERADE"
}
chain DOCKER {
iifname != "docker0" tcp dport 5100 counter packets 8247 bytes 494812 xt target "DNAT"
iifname != "docker0" tcp dport 222 counter packets 2304 bytes 137444 xt target "DNAT"
iifname != "docker0" tcp dport 20000 counter packets 1532 bytes 91584 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 25 counter packets 433 bytes 23099 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 7080 counter packets 35 bytes 1864 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 110 counter packets 195 bytes 9741 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 143 counter packets 448 bytes 25580 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 7443 counter packets 58 bytes 2868 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 465 counter packets 107 bytes 6032 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 587 counter packets 447 bytes 23492 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 993 counter packets 201 bytes 10174 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 995 counter packets 197 bytes 11180 xt target "DNAT"
iifname != "br-1ccb887b3344" tcp dport 20004 counter packets 5 bytes 300 xt target "DNAT"
iifname != "br-9fd22324ec08" tcp dport 20005 counter packets 5 bytes 300 xt target "DNAT"
iifname != "br-73641cceafc3" tcp dport 20006 counter packets 19 bytes 1140 xt target "DNAT"
iifname != "br-e5d78502832d" tcp dport 20007 counter packets 5 bytes 284 xt target "DNAT"
iifname != "br-e5d78502832d" tcp dport 20008 counter packets 4 bytes 240 xt target "DNAT"
iifname != "br-e99ce5248c84" tcp dport 1842 counter packets 12 bytes 720 xt target "DNAT"
iifname != "br-8ce143481a5b" tcp dport 4848 counter packets 40 bytes 1960 xt target "DNAT"
iifname != "docker0" tcp dport 4222 counter packets 3846 bytes 231012 xt target "DNAT"
ip daddr 127.0.0.1 iifname != "docker0" tcp dport 8222 counter packets 0 bytes 0 xt target "DNAT"
iifname != "docker0" tcp dport 20003 counter packets 18942 bytes 1136512 xt target "DNAT"
iifname != "br-2e4a76a7cd2e" tcp dport 9070 counter packets 195 bytes 11676 xt target "DNAT"
iifname != "br-77eb8a9e2ba1" tcp dport 9102 counter packets 13 bytes 772 xt target "DNAT"
iifname != "br-77eb8a9e2ba1" tcp dport 8102 counter packets 17 bytes 944 xt target "DNAT"
iifname != "br-77eb8a9e2ba1" tcp dport 8104 counter packets 16 bytes 916 xt target "DNAT"
iifname != "br-77eb8a9e2ba1" tcp dport 8103 counter packets 13 bytes 756 xt target "DNAT"
iifname != "br-3008d408e73a" tcp dport 1212 counter packets 189 bytes 11188 xt target "DNAT"
iifname != "br-6eb1e7f7f847" tcp dport 20009 counter packets 138 bytes 8280 xt target "DNAT"
iifname != "br-f8b083119d99" tcp dport 20001 counter packets 47780 bytes 2866736 xt target "DNAT"
iifname != "br-f8b083119d99" tcp dport 20002 counter packets 7 bytes 404 xt target "DNAT"
iifname != "docker0" tcp dport 20010 counter packets 74 bytes 4424 xt target "DNAT"
iifname != "docker0" tcp dport 20011 counter packets 4 bytes 240 xt target "DNAT"
iifname != "br-72fd626a8ff7" tcp dport 20012 counter packets 237 bytes 14220 xt target "DNAT"
iifname != "docker0" tcp dport 6852 counter packets 16489 bytes 989324 xt target "DNAT"
}
}
# Warning: table ip6 nat is managed by iptables-nft, do not touch!
table ip6 nat {
chain PREROUTING {
type nat hook prerouting priority dstnat; policy accept;
xt match "addrtype" counter packets 399 bytes 22104 jump DOCKER
}
chain OUTPUT {
type nat hook output priority dstnat; policy accept;
ip6 daddr != ::1 xt match "addrtype" counter packets 0 bytes 0 jump DOCKER
}
chain DOCKER {
}
}
table ip raw {
chain PREROUTING {
type filter hook prerouting priority raw; policy accept;
ip daddr 127.0.0.1 iifname != "lo" tcp dport 8222 counter packets 0 bytes 0 drop
}
}
table ip mangle {
chain FORWARD {
type filter hook forward priority mangle; policy accept;
}
}
table inet mesh {
chain input {
type filter hook input priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iif "lo" accept
iifname != { "mesh0", "enp9s0" } accept
icmp type echo-request accept
icmpv6 type { echo-request, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept
iifname != { "mesh0", "enp9s0" } udp dport { 53, 67 } accept
iifname != { "mesh0", "enp9s0" } tcp dport 53 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 22 accept
tcp dport 22 accept
tcp dport 4222 accept
tcp dport 22 accept
tcp dport 25 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 53 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } udp dport 53 accept
tcp dport 80 accept
tcp dport 110 accept
tcp dport 143 accept
tcp dport 222 accept
tcp dport 443 accept
tcp dport 465 accept
tcp dport 587 accept
tcp dport 993 accept
tcp dport 995 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 1212 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 1842 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 4222 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 4848 accept
tcp dport 5100 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 6852 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 7080 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 7443 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 8102 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 8103 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 8104 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 9000 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 9070 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 9102 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20000 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20001 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20002 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20003 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20004 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20005 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20006 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20007 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20008 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20009 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20010 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20011 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 20012 accept
udp dport 51820 accept
}
chain output {
type filter hook output priority filter; policy accept;
}
chain forward {
type filter hook forward priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iifname != { "mesh0", "enp9s0" } accept
iifname "mesh0" oifname "mesh0" accept
ct original proto-dst 22 accept
ct original proto-dst 25 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 53 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 53 accept
ct original proto-dst 80 accept
ct original proto-dst 110 accept
ct original proto-dst 143 accept
ct original proto-dst 222 accept
ct original proto-dst 443 accept
ct original proto-dst 465 accept
ct original proto-dst 587 accept
ct original proto-dst 993 accept
ct original proto-dst 995 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 1212 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 1842 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 4222 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 4848 accept
ct original proto-dst 5100 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 6852 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 7080 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 7443 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 8102 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 8103 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 8104 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 9000 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 9070 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 9102 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20000 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20001 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20002 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20003 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20004 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20005 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20006 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20007 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20008 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20009 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20010 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20011 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 20012 accept
ct original proto-dst 51820 accept
ct original proto-dst 4222 accept
}
}
+147
View File
@@ -0,0 +1,147 @@
-P INPUT ACCEPT
-P FORWARD DROP
-P OUTPUT ACCEPT
-N DOCKER
-N DOCKER-BRIDGE
-N DOCKER-CT
-N DOCKER-FORWARD
-N DOCKER-INTERNAL
-N DOCKER-USER
-N f2b-route-proxy
-N ufw-after-forward
-N ufw-after-input
-N ufw-after-logging-forward
-N ufw-after-logging-input
-N ufw-after-logging-output
-N ufw-after-output
-N ufw-before-forward
-N ufw-before-input
-N ufw-before-logging-forward
-N ufw-before-logging-input
-N ufw-before-logging-output
-N ufw-before-output
-N ufw-reject-forward
-N ufw-reject-input
-N ufw-reject-output
-N ufw-track-forward
-N ufw-track-input
-N ufw-track-output
-A INPUT -p tcp -j f2b-route-proxy
-A INPUT -j ufw-before-logging-input
-A INPUT -j ufw-before-input
-A INPUT -j ufw-after-input
-A INPUT -j ufw-after-logging-input
-A INPUT -j ufw-reject-input
-A INPUT -j ufw-track-input
-A FORWARD -j DOCKER-USER
-A FORWARD -j DOCKER-FORWARD
-A FORWARD -j ufw-before-logging-forward
-A FORWARD -j ufw-before-forward
-A FORWARD -j ufw-after-forward
-A FORWARD -j ufw-after-logging-forward
-A FORWARD -j ufw-reject-forward
-A FORWARD -j ufw-track-forward
-A OUTPUT -j ufw-before-logging-output
-A OUTPUT -j ufw-before-output
-A OUTPUT -j ufw-after-output
-A OUTPUT -j ufw-after-logging-output
-A OUTPUT -j ufw-reject-output
-A OUTPUT -j ufw-track-output
-A DOCKER -d 172.17.0.18/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8686 -j ACCEPT
-A DOCKER -d 172.17.0.14/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8989 -j ACCEPT
-A DOCKER -d 172.17.0.15/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 7878 -j ACCEPT
-A DOCKER -d 172.17.0.5/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 9117 -j ACCEPT
-A DOCKER -d 172.17.0.13/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 6789 -j ACCEPT
-A DOCKER -d 172.19.0.2/32 ! -i br-32062158f584 -o br-32062158f584 -p tcp -m tcp --dport 8080 -j ACCEPT
-A DOCKER -d 172.17.0.2/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 5432 -j ACCEPT
-A DOCKER -d 172.27.0.2/32 ! -i br-0910a98c6158 -o br-0910a98c6158 -p tcp -m tcp --dport 5678 -j ACCEPT
-A DOCKER -d 172.17.0.21/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3579 -j ACCEPT
-A DOCKER -d 172.17.0.19/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8181 -j ACCEPT
-A DOCKER -d 172.17.0.17/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8787 -j ACCEPT
-A DOCKER -d 172.17.0.16/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 6767 -j ACCEPT
-A DOCKER -d 172.17.0.12/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3000 -j ACCEPT
-A DOCKER -d 172.17.0.11/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 80 -j ACCEPT
-A DOCKER -d 172.17.0.10/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 9443 -j ACCEPT
-A DOCKER -d 172.17.0.10/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 9000 -j ACCEPT
-A DOCKER -d 172.17.0.9/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3000 -j ACCEPT
-A DOCKER -d 172.17.0.7/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 1880 -j ACCEPT
-A DOCKER -d 172.28.0.2/32 ! -i br-b11461b5b028 -o br-b11461b5b028 -p tcp -m tcp --dport 80 -j ACCEPT
-A DOCKER -d 172.23.0.14/32 ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -p tcp -m tcp --dport 6543 -j ACCEPT
-A DOCKER -d 172.23.0.14/32 ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -p tcp -m tcp --dport 5432 -j ACCEPT
-A DOCKER -d 172.23.0.5/32 ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -p tcp -m tcp --dport 8000 -j ACCEPT
-A DOCKER -d 172.26.0.3/32 ! -i br-b0fec361ccaa -o br-b0fec361ccaa -p tcp -m tcp --dport 6167 -j ACCEPT
-A DOCKER -d 172.26.0.2/32 ! -i br-b0fec361ccaa -o br-b0fec361ccaa -p tcp -m tcp --dport 80 -j ACCEPT
-A DOCKER -d 172.17.0.8/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8000 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 10001 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8880 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8843 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8443 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8080 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 6789 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 5514 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 3478 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 1900 -j ACCEPT
-A DOCKER -d 172.25.0.3/32 ! -i br-b98821f7dc38 -o br-b98821f7dc38 -p tcp -m tcp --dport 8000 -j ACCEPT
-A DOCKER -d 172.18.0.3/32 ! -i br-442a0bfc65f8 -o br-442a0bfc65f8 -p tcp -m tcp --dport 1433 -j ACCEPT
-A DOCKER -d 172.20.0.3/32 ! -i br-afa37ac8b33d -o br-afa37ac8b33d -p tcp -m tcp --dport 8081 -j ACCEPT
-A DOCKER -d 172.20.0.3/32 ! -i br-afa37ac8b33d -o br-afa37ac8b33d -p tcp -m tcp --dport 1883 -j ACCEPT
-A DOCKER -d 172.17.0.4/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8086 -j ACCEPT
-A DOCKER -d 172.21.0.2/32 ! -i br-df15d8e19ec7 -o br-df15d8e19ec7 -p tcp -m tcp --dport 6379 -j ACCEPT
-A DOCKER -d 172.30.0.3/32 ! -i br-521eab9a3a5e -o br-521eab9a3a5e -p tcp -m tcp --dport 8283 -j ACCEPT
-A DOCKER -d 172.17.0.3/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3000 -j ACCEPT
-A DOCKER ! -i br-32062158f584 -o br-32062158f584 -j DROP
-A DOCKER ! -i docker0 -o docker0 -j DROP
-A DOCKER ! -i br-521eab9a3a5e -o br-521eab9a3a5e -j DROP
-A DOCKER ! -i br-df15d8e19ec7 -o br-df15d8e19ec7 -j DROP
-A DOCKER ! -i br-afa37ac8b33d -o br-afa37ac8b33d -j DROP
-A DOCKER ! -i br-442a0bfc65f8 -o br-442a0bfc65f8 -j DROP
-A DOCKER ! -i br-b98821f7dc38 -o br-b98821f7dc38 -j DROP
-A DOCKER ! -i br-b0fec361ccaa -o br-b0fec361ccaa -j DROP
-A DOCKER ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -j DROP
-A DOCKER ! -i br-b11461b5b028 -o br-b11461b5b028 -j DROP
-A DOCKER ! -i br-2df4e541b877 -o br-2df4e541b877 -j DROP
-A DOCKER ! -i br-0910a98c6158 -o br-0910a98c6158 -j DROP
-A DOCKER-BRIDGE -o br-32062158f584 -j DOCKER
-A DOCKER-BRIDGE -o docker0 -j DOCKER
-A DOCKER-BRIDGE -o br-521eab9a3a5e -j DOCKER
-A DOCKER-BRIDGE -o br-df15d8e19ec7 -j DOCKER
-A DOCKER-BRIDGE -o br-afa37ac8b33d -j DOCKER
-A DOCKER-BRIDGE -o br-442a0bfc65f8 -j DOCKER
-A DOCKER-BRIDGE -o br-b98821f7dc38 -j DOCKER
-A DOCKER-BRIDGE -o br-b0fec361ccaa -j DOCKER
-A DOCKER-BRIDGE -o br-66ffa5c1cba5 -j DOCKER
-A DOCKER-BRIDGE -o br-b11461b5b028 -j DOCKER
-A DOCKER-BRIDGE -o br-2df4e541b877 -j DOCKER
-A DOCKER-BRIDGE -o br-0910a98c6158 -j DOCKER
-A DOCKER-CT -o br-32062158f584 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o docker0 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-521eab9a3a5e -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-df15d8e19ec7 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-afa37ac8b33d -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-442a0bfc65f8 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-b98821f7dc38 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-b0fec361ccaa -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-66ffa5c1cba5 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-b11461b5b028 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-2df4e541b877 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-0910a98c6158 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-FORWARD -j DOCKER-CT
-A DOCKER-FORWARD -j DOCKER-INTERNAL
-A DOCKER-FORWARD -j DOCKER-BRIDGE
-A DOCKER-FORWARD -i br-32062158f584 -j ACCEPT
-A DOCKER-FORWARD -i docker0 -j ACCEPT
-A DOCKER-FORWARD -i br-521eab9a3a5e -j ACCEPT
-A DOCKER-FORWARD -i br-df15d8e19ec7 -j ACCEPT
-A DOCKER-FORWARD -i br-afa37ac8b33d -j ACCEPT
-A DOCKER-FORWARD -i br-442a0bfc65f8 -j ACCEPT
-A DOCKER-FORWARD -i br-b98821f7dc38 -j ACCEPT
-A DOCKER-FORWARD -i br-b0fec361ccaa -j ACCEPT
-A DOCKER-FORWARD -i br-66ffa5c1cba5 -j ACCEPT
-A DOCKER-FORWARD -i br-b11461b5b028 -j ACCEPT
-A DOCKER-FORWARD -i br-2df4e541b877 -j ACCEPT
-A DOCKER-FORWARD -i br-0910a98c6158 -j ACCEPT
-A DOCKER-USER -p tcp -j f2b-route-proxy
-A f2b-route-proxy -s 13.70.107.184/32 -j REJECT --reject-with icmp-port-unreachable
-A f2b-route-proxy -s 45.138.12.51/32 -j REJECT --reject-with icmp-port-unreachable
-A f2b-route-proxy -s 20.214.191.94/32 -j REJECT --reject-with icmp-port-unreachable
-A f2b-route-proxy -j RETURN
+149
View File
@@ -0,0 +1,149 @@
-P INPUT ACCEPT
-P FORWARD DROP
-P OUTPUT ACCEPT
-N DOCKER
-N DOCKER-BRIDGE
-N DOCKER-CT
-N DOCKER-FORWARD
-N DOCKER-INTERNAL
-N DOCKER-USER
-N HAL-MESH-ONLY
-N ufw-after-forward
-N ufw-after-input
-N ufw-after-logging-forward
-N ufw-after-logging-input
-N ufw-after-logging-output
-N ufw-after-output
-N ufw-before-forward
-N ufw-before-input
-N ufw-before-logging-forward
-N ufw-before-logging-input
-N ufw-before-logging-output
-N ufw-before-output
-N ufw-reject-forward
-N ufw-reject-input
-N ufw-reject-output
-N ufw-track-forward
-N ufw-track-input
-N ufw-track-output
-A INPUT -j ufw-before-logging-input
-A INPUT -j ufw-before-input
-A INPUT -j ufw-after-input
-A INPUT -j ufw-after-logging-input
-A INPUT -j ufw-reject-input
-A INPUT -j ufw-track-input
-A FORWARD -j DOCKER-USER
-A FORWARD -j DOCKER-FORWARD
-A FORWARD -j ufw-before-logging-forward
-A FORWARD -j ufw-before-forward
-A FORWARD -j ufw-after-forward
-A FORWARD -j ufw-after-logging-forward
-A FORWARD -j ufw-reject-forward
-A FORWARD -j ufw-track-forward
-A OUTPUT -j ufw-before-logging-output
-A OUTPUT -j ufw-before-output
-A OUTPUT -j ufw-after-output
-A OUTPUT -j ufw-after-logging-output
-A OUTPUT -j ufw-reject-output
-A OUTPUT -j ufw-track-output
-A DOCKER -d 172.17.0.18/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8686 -j ACCEPT
-A DOCKER -d 172.17.0.14/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8989 -j ACCEPT
-A DOCKER -d 172.17.0.15/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 7878 -j ACCEPT
-A DOCKER -d 172.17.0.5/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 9117 -j ACCEPT
-A DOCKER -d 172.17.0.13/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 6789 -j ACCEPT
-A DOCKER -d 172.19.0.2/32 ! -i br-32062158f584 -o br-32062158f584 -p tcp -m tcp --dport 8080 -j ACCEPT
-A DOCKER -d 172.17.0.2/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 5432 -j ACCEPT
-A DOCKER -d 172.27.0.2/32 ! -i br-0910a98c6158 -o br-0910a98c6158 -p tcp -m tcp --dport 5678 -j ACCEPT
-A DOCKER -d 172.17.0.21/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3579 -j ACCEPT
-A DOCKER -d 172.17.0.19/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8181 -j ACCEPT
-A DOCKER -d 172.17.0.17/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8787 -j ACCEPT
-A DOCKER -d 172.17.0.16/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 6767 -j ACCEPT
-A DOCKER -d 172.17.0.12/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3000 -j ACCEPT
-A DOCKER -d 172.17.0.11/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 80 -j ACCEPT
-A DOCKER -d 172.17.0.10/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 9443 -j ACCEPT
-A DOCKER -d 172.17.0.10/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 9000 -j ACCEPT
-A DOCKER -d 172.17.0.9/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3000 -j ACCEPT
-A DOCKER -d 172.17.0.7/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 1880 -j ACCEPT
-A DOCKER -d 172.28.0.2/32 ! -i br-b11461b5b028 -o br-b11461b5b028 -p tcp -m tcp --dport 80 -j ACCEPT
-A DOCKER -d 172.23.0.14/32 ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -p tcp -m tcp --dport 6543 -j ACCEPT
-A DOCKER -d 172.23.0.14/32 ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -p tcp -m tcp --dport 5432 -j ACCEPT
-A DOCKER -d 172.23.0.5/32 ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -p tcp -m tcp --dport 8000 -j ACCEPT
-A DOCKER -d 172.26.0.3/32 ! -i br-b0fec361ccaa -o br-b0fec361ccaa -p tcp -m tcp --dport 6167 -j ACCEPT
-A DOCKER -d 172.26.0.2/32 ! -i br-b0fec361ccaa -o br-b0fec361ccaa -p tcp -m tcp --dport 80 -j ACCEPT
-A DOCKER -d 172.17.0.8/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8000 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 10001 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8880 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8843 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8443 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8080 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 6789 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 5514 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 3478 -j ACCEPT
-A DOCKER -d 172.17.0.6/32 ! -i docker0 -o docker0 -p udp -m udp --dport 1900 -j ACCEPT
-A DOCKER -d 172.25.0.3/32 ! -i br-b98821f7dc38 -o br-b98821f7dc38 -p tcp -m tcp --dport 8000 -j ACCEPT
-A DOCKER -d 172.18.0.3/32 ! -i br-442a0bfc65f8 -o br-442a0bfc65f8 -p tcp -m tcp --dport 1433 -j ACCEPT
-A DOCKER -d 172.20.0.3/32 ! -i br-afa37ac8b33d -o br-afa37ac8b33d -p tcp -m tcp --dport 8081 -j ACCEPT
-A DOCKER -d 172.20.0.3/32 ! -i br-afa37ac8b33d -o br-afa37ac8b33d -p tcp -m tcp --dport 1883 -j ACCEPT
-A DOCKER -d 172.17.0.4/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 8086 -j ACCEPT
-A DOCKER -d 172.21.0.2/32 ! -i br-df15d8e19ec7 -o br-df15d8e19ec7 -p tcp -m tcp --dport 6379 -j ACCEPT
-A DOCKER -d 172.30.0.3/32 ! -i br-521eab9a3a5e -o br-521eab9a3a5e -p tcp -m tcp --dport 8283 -j ACCEPT
-A DOCKER -d 172.17.0.3/32 ! -i docker0 -o docker0 -p tcp -m tcp --dport 3000 -j ACCEPT
-A DOCKER ! -i br-32062158f584 -o br-32062158f584 -j DROP
-A DOCKER ! -i docker0 -o docker0 -j DROP
-A DOCKER ! -i br-521eab9a3a5e -o br-521eab9a3a5e -j DROP
-A DOCKER ! -i br-df15d8e19ec7 -o br-df15d8e19ec7 -j DROP
-A DOCKER ! -i br-afa37ac8b33d -o br-afa37ac8b33d -j DROP
-A DOCKER ! -i br-442a0bfc65f8 -o br-442a0bfc65f8 -j DROP
-A DOCKER ! -i br-b98821f7dc38 -o br-b98821f7dc38 -j DROP
-A DOCKER ! -i br-b0fec361ccaa -o br-b0fec361ccaa -j DROP
-A DOCKER ! -i br-66ffa5c1cba5 -o br-66ffa5c1cba5 -j DROP
-A DOCKER ! -i br-b11461b5b028 -o br-b11461b5b028 -j DROP
-A DOCKER ! -i br-2df4e541b877 -o br-2df4e541b877 -j DROP
-A DOCKER ! -i br-0910a98c6158 -o br-0910a98c6158 -j DROP
-A DOCKER-BRIDGE -o br-32062158f584 -j DOCKER
-A DOCKER-BRIDGE -o docker0 -j DOCKER
-A DOCKER-BRIDGE -o br-521eab9a3a5e -j DOCKER
-A DOCKER-BRIDGE -o br-df15d8e19ec7 -j DOCKER
-A DOCKER-BRIDGE -o br-afa37ac8b33d -j DOCKER
-A DOCKER-BRIDGE -o br-442a0bfc65f8 -j DOCKER
-A DOCKER-BRIDGE -o br-b98821f7dc38 -j DOCKER
-A DOCKER-BRIDGE -o br-b0fec361ccaa -j DOCKER
-A DOCKER-BRIDGE -o br-66ffa5c1cba5 -j DOCKER
-A DOCKER-BRIDGE -o br-b11461b5b028 -j DOCKER
-A DOCKER-BRIDGE -o br-2df4e541b877 -j DOCKER
-A DOCKER-BRIDGE -o br-0910a98c6158 -j DOCKER
-A DOCKER-CT -o br-32062158f584 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o docker0 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-521eab9a3a5e -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-df15d8e19ec7 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-afa37ac8b33d -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-442a0bfc65f8 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-b98821f7dc38 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-b0fec361ccaa -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-66ffa5c1cba5 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-b11461b5b028 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-2df4e541b877 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-CT -o br-0910a98c6158 -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-FORWARD -j DOCKER-CT
-A DOCKER-FORWARD -j DOCKER-INTERNAL
-A DOCKER-FORWARD -j DOCKER-BRIDGE
-A DOCKER-FORWARD -i br-32062158f584 -j ACCEPT
-A DOCKER-FORWARD -i docker0 -j ACCEPT
-A DOCKER-FORWARD -i br-521eab9a3a5e -j ACCEPT
-A DOCKER-FORWARD -i br-df15d8e19ec7 -j ACCEPT
-A DOCKER-FORWARD -i br-afa37ac8b33d -j ACCEPT
-A DOCKER-FORWARD -i br-442a0bfc65f8 -j ACCEPT
-A DOCKER-FORWARD -i br-b98821f7dc38 -j ACCEPT
-A DOCKER-FORWARD -i br-b0fec361ccaa -j ACCEPT
-A DOCKER-FORWARD -i br-66ffa5c1cba5 -j ACCEPT
-A DOCKER-FORWARD -i br-b11461b5b028 -j ACCEPT
-A DOCKER-FORWARD -i br-2df4e541b877 -j ACCEPT
-A DOCKER-FORWARD -i br-0910a98c6158 -j ACCEPT
-A DOCKER-USER -i enp6s0 -p tcp -m conntrack --ctstate NEW -j HAL-MESH-ONLY
-A HAL-MESH-ONLY -m conntrack --ctorigdstport 6881 -j RETURN
-A HAL-MESH-ONLY -m conntrack --ctorigdstport 80 -j RETURN
-A HAL-MESH-ONLY -m conntrack --ctorigdstport 443 -j RETURN
-A HAL-MESH-ONLY -s 10.0.0.0/8 -j RETURN
-A HAL-MESH-ONLY -s 172.16.0.0/12 -j RETURN
-A HAL-MESH-ONLY -s 192.168.0.0/16 -j RETURN
-A HAL-MESH-ONLY -m comment --comment "HAL: not public -> mesh only" -j DROP
+327
View File
@@ -0,0 +1,327 @@
table ip mangle {
chain FORWARD {
type filter hook forward priority mangle; policy accept;
}
}
# Warning: table ip nat is managed by iptables-nft, do not touch!
table ip nat {
chain PREROUTING {
type nat hook prerouting priority dstnat; policy accept;
xt match "addrtype" counter packets 12072 bytes 4564241 jump DOCKER
}
chain OUTPUT {
type nat hook output priority dstnat; policy accept;
ip daddr != 127.0.0.0/8 xt match "addrtype" counter packets 1854 bytes 111240 jump DOCKER
}
chain POSTROUTING {
type nat hook postrouting priority srcnat; policy accept;
ip saddr 172.17.0.0/16 oifname != "docker0" counter packets 903 bytes 61577 xt target "MASQUERADE"
ip saddr 172.21.0.0/16 oifname != "br-86a5d6b30e2b" counter packets 344 bytes 27744 xt target "MASQUERADE"
ip saddr 172.25.0.0/16 oifname != "br-61495e14a004" counter packets 374 bytes 33016 xt target "MASQUERADE"
ip saddr 172.30.0.0/16 oifname != "br-5107796ee9b4" counter packets 352 bytes 28224 xt target "MASQUERADE"
ip saddr 172.18.0.0/16 oifname != "br-cfd337ac4e58" counter packets 339 bytes 27444 xt target "MASQUERADE"
ip saddr 172.19.0.0/16 oifname != "br-8f0c6ee01425" counter packets 351 bytes 28164 xt target "MASQUERADE"
ip saddr 172.22.0.0/16 oifname != "br-75ac3c36e87f" counter packets 333 bytes 27084 xt target "MASQUERADE"
ip saddr 172.20.0.0/16 oifname != "br-0529801521bc" counter packets 343 bytes 27404 xt target "MASQUERADE"
}
chain DOCKER {
iifname != "br-61495e14a004" tcp dport 5680 counter packets 2 bytes 120 xt target "DNAT"
ip daddr 127.0.0.1 iifname != "br-61495e14a004" tcp dport 15673 counter packets 0 bytes 0 xt target "DNAT"
ip daddr 127.0.0.1 iifname != "docker0" tcp dport 55432 counter packets 0 bytes 0 xt target "DNAT"
}
}
# Warning: table ip filter is managed by iptables-nft, do not touch!
table ip filter {
chain DOCKER-FORWARD {
counter packets 702328 bytes 1801910999 jump DOCKER-CT
counter packets 337315 bytes 23928864 jump DOCKER-INTERNAL
counter packets 337315 bytes 23928864 jump DOCKER-BRIDGE
iifname "br-75ac3c36e87f" counter packets 0 bytes 0 accept
iifname "br-86a5d6b30e2b" counter packets 0 bytes 0 accept
iifname "br-8f0c6ee01425" counter packets 0 bytes 0 accept
iifname "br-cfd337ac4e58" counter packets 0 bytes 0 accept
iifname "br-0529801521bc" counter packets 0 bytes 0 accept
iifname "br-5107796ee9b4" counter packets 0 bytes 0 accept
iifname "br-61495e14a004" counter packets 0 bytes 0 accept
iifname "docker0" counter packets 337315 bytes 23928864 accept
}
chain FORWARD {
type filter hook forward priority filter; policy drop;
counter packets 702328 bytes 1801910999 jump DOCKER-USER
counter packets 702328 bytes 1801910999 jump DOCKER-FORWARD
}
chain DOCKER-USER {
ip protocol tcp counter packets 702510 bytes 1801965868 jump f2b-sshd
oifname "incusbr0" counter packets 0 bytes 0 accept
iifname "incusbr0" counter packets 0 bytes 0 accept
}
chain f2b-sshd {
counter packets 10423854 bytes 13891318049 return
}
chain INPUT {
type filter hook input priority filter; policy accept;
ip protocol tcp counter packets 9721344 bytes 12089352181 jump f2b-sshd
}
chain DOCKER {
ip daddr 172.17.0.2 iifname != "docker0" oifname "docker0" tcp dport 5432 counter packets 0 bytes 0 accept
ip daddr 172.25.0.2 iifname != "br-61495e14a004" oifname "br-61495e14a004" tcp dport 15672 counter packets 0 bytes 0 accept
ip daddr 172.25.0.2 iifname != "br-61495e14a004" oifname "br-61495e14a004" tcp dport 5672 counter packets 0 bytes 0 accept
iifname != "br-75ac3c36e87f" oifname "br-75ac3c36e87f" counter packets 0 bytes 0 drop
iifname != "br-86a5d6b30e2b" oifname "br-86a5d6b30e2b" counter packets 0 bytes 0 drop
iifname != "br-8f0c6ee01425" oifname "br-8f0c6ee01425" counter packets 0 bytes 0 drop
iifname != "br-cfd337ac4e58" oifname "br-cfd337ac4e58" counter packets 0 bytes 0 drop
iifname != "br-0529801521bc" oifname "br-0529801521bc" counter packets 0 bytes 0 drop
iifname != "br-5107796ee9b4" oifname "br-5107796ee9b4" counter packets 0 bytes 0 drop
iifname != "br-61495e14a004" oifname "br-61495e14a004" counter packets 0 bytes 0 drop
iifname != "docker0" oifname "docker0" counter packets 0 bytes 0 drop
}
chain DOCKER-BRIDGE {
oifname "br-75ac3c36e87f" counter packets 0 bytes 0 jump DOCKER
oifname "br-86a5d6b30e2b" counter packets 0 bytes 0 jump DOCKER
oifname "br-8f0c6ee01425" counter packets 0 bytes 0 jump DOCKER
oifname "br-cfd337ac4e58" counter packets 0 bytes 0 jump DOCKER
oifname "br-0529801521bc" counter packets 0 bytes 0 jump DOCKER
oifname "br-5107796ee9b4" counter packets 0 bytes 0 jump DOCKER
oifname "br-61495e14a004" counter packets 0 bytes 0 jump DOCKER
oifname "docker0" counter packets 0 bytes 0 jump DOCKER
}
chain DOCKER-CT {
oifname "br-75ac3c36e87f" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-86a5d6b30e2b" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-8f0c6ee01425" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-cfd337ac4e58" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-0529801521bc" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-5107796ee9b4" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "br-61495e14a004" xt match "conntrack" counter packets 0 bytes 0 accept
oifname "docker0" xt match "conntrack" counter packets 365013 bytes 1777982135 accept
}
chain DOCKER-INTERNAL {
}
}
# Warning: table ip6 nat is managed by iptables-nft, do not touch!
table ip6 nat {
chain PREROUTING {
type nat hook prerouting priority dstnat; policy accept;
xt match "addrtype" counter packets 363 bytes 67927 jump DOCKER
}
chain OUTPUT {
type nat hook output priority dstnat; policy accept;
ip6 daddr != ::1 xt match "addrtype" counter packets 0 bytes 0 jump DOCKER
}
chain DOCKER {
}
}
table ip6 filter {
chain DOCKER-FORWARD {
counter packets 0 bytes 0 jump DOCKER-CT
counter packets 0 bytes 0 jump DOCKER-INTERNAL
counter packets 0 bytes 0 jump DOCKER-BRIDGE
}
chain FORWARD {
type filter hook forward priority filter; policy accept;
counter packets 0 bytes 0 jump DOCKER-USER
counter packets 0 bytes 0 jump DOCKER-FORWARD
}
chain DOCKER-USER {
}
chain DOCKER {
}
chain DOCKER-BRIDGE {
}
chain DOCKER-CT {
}
chain DOCKER-INTERNAL {
}
}
table ip raw {
chain PREROUTING {
type filter hook prerouting priority raw; policy accept;
ip daddr 172.25.0.2 iifname != "br-61495e14a004" counter packets 0 bytes 0 drop
ip daddr 127.0.0.1 iifname != "lo" tcp dport 15673 counter packets 0 bytes 0 drop
ip daddr 172.17.0.2 iifname != "docker0" counter packets 0 bytes 0 drop
ip daddr 127.0.0.1 iifname != "lo" tcp dport 55432 counter packets 0 bytes 0 drop
}
}
table inet incus {
set bridges {
type ifname
elements = { "incusbr0" }
}
chain pstrt.incusbr0 {
type nat hook postrouting priority srcnat; policy accept;
ip saddr 10.7.169.0/24 oifname @bridges accept
ip saddr 10.7.169.0/24 ip daddr != 10.7.169.0/24 masquerade
ip6 saddr fd42:cbc4:e123:f6::/64 oifname @bridges accept
ip6 saddr fd42:cbc4:e123:f6::/64 ip6 daddr != fd42:cbc4:e123:f6::/64 masquerade
}
chain fwd.incusbr0 {
type filter hook forward priority filter; policy accept;
ip version 4 oifname "incusbr0" accept
ip version 4 iifname "incusbr0" accept
ip6 version 6 oifname "incusbr0" accept
ip6 version 6 iifname "incusbr0" accept
}
chain in.incusbr0 {
type filter hook input priority filter; policy accept;
iifname "incusbr0" tcp dport 53 accept
iifname "incusbr0" udp dport 53 accept
iifname "incusbr0" icmp type { destination-unreachable, time-exceeded, parameter-problem } accept
iifname "incusbr0" udp dport 67 accept
iifname "incusbr0" ip protocol udp udp checksum set 0
iifname "incusbr0" icmpv6 type { destination-unreachable, packet-too-big, time-exceeded, parameter-problem, nd-router-solicit, nd-neighbor-solicit, nd-neighbor-advert, mld2-listener-report } accept
iifname "incusbr0" udp dport 547 accept
}
chain out.incusbr0 {
type filter hook output priority filter; policy accept;
oifname "incusbr0" tcp sport 53 accept
oifname "incusbr0" udp sport 53 accept
oifname "incusbr0" icmp type { destination-unreachable, time-exceeded, parameter-problem } accept
oifname "incusbr0" udp sport 67 accept
oifname "incusbr0" ip protocol udp udp checksum set 0
oifname "incusbr0" icmpv6 type { destination-unreachable, packet-too-big, time-exceeded, parameter-problem, echo-request, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert, mld2-listener-report } accept
oifname "incusbr0" udp sport 547 accept
}
}
table ip fct_filter {
chain OUTPUT {
type filter hook output priority filter; policy accept;
}
chain FCT-QUARANTINE-EMS {
}
chain FCT-QUARANTINE-FAZ {
}
chain FORWARD {
type filter hook forward priority filter; policy accept;
}
chain FCT-WEBFILTER-QUIC-CHAIN {
}
chain INPUT {
type filter hook input priority filter; policy accept;
}
chain FCT-QUARANTINE {
}
chain FCT-DNS-QUIC-FILTER {
}
chain FCT-VPN-CHAIN {
}
}
table ip fct_nat {
chain OUTPUT {
type nat hook output priority dstnat; policy accept;
}
chain FCT-DNS-UDP-CHAIN-STAGE-2 {
}
chain FCT-DNS-UDP-CHAIN-STAGE-1 {
}
chain FCT-TCP-CHAIN {
}
chain FCT-DNS-DOH-CHAIN-STAGE-1 {
}
chain FCT-WEBFILTER-CHAIN {
}
chain FCT-DNS-DOH-CHAIN-STAGE-2 {
}
}
table ip6 fct_filter {
chain FCT-QUARANTINE {
}
chain INPUT {
type filter hook input priority filter; policy accept;
}
chain FORWARD {
type filter hook forward priority filter; policy accept;
}
chain OUTPUT {
type filter hook output priority filter; policy accept;
}
}
table ip fct_mangle {
chain PREROUTING {
type filter hook prerouting priority mangle; policy accept;
}
chain FCT-UDP-STAGE-1 {
}
chain OUTPUT {
type route hook output priority mangle; policy accept;
}
chain FCT-UDP-STAGE-2 {
}
chain FCT-UDP-OUTPUT {
}
}
table inet mesh {
chain input {
type filter hook input priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iif "lo" accept
iifname != { "mesh0", "wlp3s0" } accept
icmp type echo-request accept
icmpv6 type { echo-request, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept
iifname != { "mesh0", "wlp3s0" } udp dport { 53, 67 } accept
iifname != { "mesh0", "wlp3s0" } tcp dport 53 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 22 accept
tcp dport 22 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } tcp dport 53 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } udp dport 53 accept
}
chain output {
type filter hook output priority filter; policy accept;
}
chain forward {
type filter hook forward priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iifname != { "mesh0", "wlp3s0" } accept
iifname "mesh0" oifname "mesh0" accept
ct original proto-dst 22 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 53 accept
ip saddr { 10.10.0.1, 10.10.0.2, 10.10.0.3, 10.10.0.4 } ct original proto-dst 53 accept
}
}
+26
View File
@@ -110,6 +110,17 @@ type Report struct {
// and the mesh's up in its place, and where the found configuration's original was kept. // and the mesh's up in its place, and where the found configuration's original was kept.
Tunnel *CarriedTunnel `json:"tunnel,omitempty"` Tunnel *CarriedTunnel `json:"tunnel,omitempty"`
// Filters is what filters this machine now, every table and chain that refuses traffic with its
// owner — the mesh's, the found firewall's, the container runtime's own, a ban list, or other
// (novox/hq ADR 0168). Every node reports it, adopted or converged, so the mesh can say
// truthfully what filters a converged machine and name what it did not write.
Filters []Filter `json:"filters,omitempty"`
// FoundFirewall is the state of the firewall a converged machine was found with: whether it is
// in force now, and how it came to be inactive — the mesh disabled it, or a reconcile found it so
// (ADR 0168). Nil on a machine found with none, and on an adopted one, where Firewall says it.
FoundFirewall *FoundFirewall `json:"found_firewall,omitempty"`
// Strays is what runs on the machine that the mesh neither wrote nor holds (novox/hq ADR // Strays is what runs on the machine that the mesh neither wrote nor holds (novox/hq ADR
// 0163): containers nobody declared and nobody holds, the ones a cutover leaves behind. // 0163): containers nobody declared and nobody holds, the ones a cutover leaves behind.
Strays []Stray `json:"strays,omitempty"` Strays []Stray `json:"strays,omitempty"`
@@ -233,3 +244,18 @@ type Reach struct {
Published bool `json:"published,omitempty"` Published bool `json:"published,omitempty"`
ContainerPort int `json:"container-port,omitempty"` ContainerPort int `json:"container-port,omitempty"`
} }
// A Filter is one place on the machine that refuses traffic, with its owner (novox/hq ADR 0168):
// the same shape the host's firewall package reads, carried as data.
type Filter struct {
Where string `json:"where"`
Owner string `json:"owner"`
Refuses string `json:"refuses"`
}
// FoundFirewall is the state of a converged machine's found firewall (ADR 0168).
type FoundFirewall struct {
Kind string `json:"kind"`
Active bool `json:"active"`
RetiredBy string `json:"retired_by,omitempty"`
}
+42 -4
View File
@@ -29,17 +29,32 @@ import (
// routing table without one. // routing table without one.
const ProcNet = "/proc/net" const ProcNet = "/proc/net"
// Links are the interfaces carrying a default route, for both address families, sorted and without // SysClassNet is where the kernel lists the machine's network interfaces, one directory each. A
// repeats. // parameter for the same reason.
const SysClassNet = "/sys/class/net"
// Links are the interfaces carrying a default route, for both address families, and every interface
// backed by a physical device, sorted and without repeats.
//
// **A physical link faces outside whether or not it is up** (novox/hq issue 197). The filter accepts
// whatever did not arrive on a link named here, so a link left out of this list is not filtered at
// all. A cable unplugged when the machine last reported carries no default route, and was left out:
// plugged in, everything arriving on it was accepted until the next report and the next push — and a
// second physical link that never carries the default route was never filtered. A physical device is
// read from the kernel's own list, where it has a `device` entry; a bridge, a veth, the tunnel and the
// loopback have none, and stay what they are, this machine's own.
// //
// A machine may have more than one: a laptop with a cable and a radio has two, and both face // A machine may have more than one: a laptop with a cable and a radio has two, and both face
// outside. A machine with none — no route off itself — returns nothing, and the mesh refuses to // outside. A machine with none — no route off itself — returns nothing, and the mesh refuses to
// compose a filter for it rather than writing a rule around a link with no name, which would be a // compose a filter for it rather than writing a rule around a link with no name, which would be a
// rule set that does not load and a machine filtering nothing while its unit reports success. // rule set that does not load and a machine filtering nothing while its unit reports success.
func Links(procNet string) ([]string, error) { func Links(procNet, sysClassNet string) ([]string, error) {
if procNet == "" { if procNet == "" {
procNet = ProcNet procNet = ProcNet
} }
if sysClassNet == "" {
sysClassNet = SysClassNet
}
seen := map[string]bool{} seen := map[string]bool{}
four, err := defaultsV4(filepath.Join(procNet, "route")) four, err := defaultsV4(filepath.Join(procNet, "route"))
@@ -50,7 +65,11 @@ func Links(procNet string) ([]string, error) {
if err != nil { if err != nil {
return nil, err return nil, err
} }
for _, name := range append(four, six...) { devices, err := physical(sysClassNet)
if err != nil {
return nil, err
}
for _, name := range append(append(four, six...), devices...) {
if name != "" && name != "lo" { if name != "" && name != "lo" {
seen[name] = true seen[name] = true
} }
@@ -64,6 +83,25 @@ func Links(procNet string) ([]string, error) {
return out, nil return out, nil
} }
// physical is every interface the kernel lists with a device behind it. A list that is not there is
// not an error — a machine without sysfs mounted reports what its routing table says, as before.
func physical(sysClassNet string) ([]string, error) {
entries, err := os.ReadDir(sysClassNet)
if os.IsNotExist(err) {
return nil, nil
}
if err != nil {
return nil, err
}
var out []string
for _, e := range entries {
if _, err := os.Stat(filepath.Join(sysClassNet, e.Name(), "device")); err == nil {
out = append(out, e.Name())
}
}
return out, nil
}
// defaultsV4 reads /proc/net/route, whose columns are // defaultsV4 reads /proc/net/route, whose columns are
// //
// Iface Destination Gateway Flags RefCnt Use Metric Mask ... // Iface Destination Gateway Flags RefCnt Use Metric Mask ...
+42 -6
View File
@@ -32,7 +32,7 @@ func TestLinksAreTheOnesCarryingADefaultRoute(t *testing.T) {
write(t, dir, "route", routeV4) write(t, dir, "route", routeV4)
write(t, dir, "ipv6_route", routeV6) write(t, dir, "ipv6_route", routeV6)
got, err := Links(dir) got, err := Links(dir, t.TempDir())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -52,7 +52,7 @@ func TestAZeroDestinationWithAMaskIsNotADefaultRoute(t *testing.T) {
write(t, dir, "route", `Iface Destination Gateway Flags RefCnt Use Metric Mask MTU Window IRTT write(t, dir, "route", `Iface Destination Gateway Flags RefCnt Use Metric Mask MTU Window IRTT
br-abc 00000000 00000000 0001 0 0 0 00FFFFFF 0 0 0 br-abc 00000000 00000000 0001 0 0 0 00FFFFFF 0 0 0
`) `)
got, err := Links(dir) got, err := Links(dir, t.TempDir())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -67,7 +67,7 @@ br-abc 00000000 00000000 0001 0 0 0 00FFFFFF 0 0 0
func TestNoDefaultRouteIsNoLinks(t *testing.T) { func TestNoDefaultRouteIsNoLinks(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
write(t, dir, "route", "Iface\tDestination\tGateway \tFlags\tRefCnt\tUse\tMetric\tMask\t\tMTU\tWindow\tIRTT\n") write(t, dir, "route", "Iface\tDestination\tGateway \tFlags\tRefCnt\tUse\tMetric\tMask\t\tMTU\tWindow\tIRTT\n")
got, err := Links(dir) got, err := Links(dir, t.TempDir())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -81,7 +81,7 @@ func TestNoDefaultRouteIsNoLinks(t *testing.T) {
func TestAMissingTableIsNotAFailure(t *testing.T) { func TestAMissingTableIsNotAFailure(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
write(t, dir, "route", routeV4) write(t, dir, "route", routeV4)
got, err := Links(dir) got, err := Links(dir, t.TempDir())
if err != nil { if err != nil {
t.Fatalf("a missing v6 table should not fail: %v", err) t.Fatalf("a missing v6 table should not fail: %v", err)
} }
@@ -97,7 +97,7 @@ func TestALinkIsReportedOnce(t *testing.T) {
write(t, dir, "ipv6_route", write(t, dir, "ipv6_route",
"00000000000000000000000000000000 00 00000000000000000000000000000000 00 "+ "00000000000000000000000000000000 00 00000000000000000000000000000000 00 "+
"fe800000000000000000000000000001 00000400 00000001 00000000 00000003 enp9s0\n") "fe800000000000000000000000000001 00000400 00000001 00000000 00000003 enp9s0\n")
got, err := Links(dir) got, err := Links(dir, t.TempDir())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -109,7 +109,7 @@ func TestALinkIsReportedOnce(t *testing.T) {
// Against this machine's own routing table, so the parse is held to what the kernel actually writes // Against this machine's own routing table, so the parse is held to what the kernel actually writes
// and not only to a fixture written to agree with it. // and not only to a fixture written to agree with it.
func TestAgainstThisMachinesOwnTable(t *testing.T) { func TestAgainstThisMachinesOwnTable(t *testing.T) {
got, err := Links("") got, err := Links("", "")
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -118,3 +118,39 @@ func TestAgainstThisMachinesOwnTable(t *testing.T) {
} }
t.Logf("this machine's outward links: %v", got) t.Logf("this machine's outward links: %v", got)
} }
// sysNet is a /sys/class/net: each name a directory, with a `device` entry when a device backs it.
func sysNet(t *testing.T, physical []string, virtual []string) string {
t.Helper()
dir := t.TempDir()
for _, name := range physical {
if err := os.MkdirAll(filepath.Join(dir, name, "device"), 0o755); err != nil {
t.Fatal(err)
}
}
for _, name := range virtual {
if err := os.MkdirAll(filepath.Join(dir, name), 0o755); err != nil {
t.Fatal(err)
}
}
return dir
}
// **A physical link faces outside whether or not it carries the default route** (novox/hq issue
// 197). A machine on its radio with its cable unplugged reported only the radio, and the filter then
// accepted everything arriving on the cable the moment it was plugged in. Bridges, veths, the tunnel
// and the loopback have no device behind them and stay this machine's own.
func TestEveryPhysicalLinkFacesOutsideUpOrDown(t *testing.T) {
proc := t.TempDir()
write(t, proc, "route", `Iface Destination Gateway Flags RefCnt Use Metric Mask MTU Window IRTT
wlp5s0 00000000 01FEA8C0 0003 0 0 600 00000000 0 0 0
`)
sys := sysNet(t, []string{"wlp5s0", "enp6s0"}, []string{"lo", "docker0", "br-0123456789ab", "veth1", "mesh0"})
got, err := Links(proc, sys)
if err != nil {
t.Fatal(err)
}
if want := []string{"enp6s0", "wlp5s0"}; !reflect.DeepEqual(got, want) {
t.Fatalf("outward links are %v, want %v", got, want)
}
}
+8
View File
@@ -15,6 +15,9 @@ const (
CapServiceManager = "service-manager" CapServiceManager = "service-manager"
CapFirewall = "firewall" CapFirewall = "firewall"
CapOverlay = "overlay" CapOverlay = "overlay"
// CapVirtualisation is a running virtualisation daemon: what the lab raises its machines on
// (novox/hq ADR 0172), and what grants a module the daemon's socket.
CapVirtualisation = "virtualisation"
CapGraphicalSession = "graphical-session" CapGraphicalSession = "graphical-session"
// CapSeat is hardware: somewhere a display server COULD run. CapGraphicalSession above is // CapSeat is hardware: somewhere a display server COULD run. CapGraphicalSession above is
// state: whether one IS running. Assignment needs the first. // state: whether one IS running. Assignment needs the first.
@@ -205,6 +208,11 @@ func Default(runner Runner) []Detector {
why: "lists the ruleset — needs the tool AND the privilege to use it", why: "lists the ruleset — needs the tool AND the privilege to use it",
runner: runner, runner: runner,
}, },
commandCapability{
name: CapVirtualisation, command: "incus", args: []string{"info"},
why: "asks the virtualisation daemon about itself — a running daemon, not an installed client",
runner: runner,
},
commandCapability{ commandCapability{
name: CapOverlay, command: "wg", args: []string{"show", "interfaces"}, name: CapOverlay, command: "wg", args: []string{"show", "interfaces"},
why: "asks the kernel for interfaces — needs the module, not just the tool", why: "asks the kernel for interfaces — needs the module, not just the tool",
+75 -1
View File
@@ -77,11 +77,21 @@ type Applied struct {
// (novox/hq issue 128) is given back its original with the mesh's region in it — and a path // (novox/hq issue 128) is given back its original with the mesh's region in it — and a path
// said once in a log line is not a path the host can find again. // said once in a log line is not a path the host can find again.
Kept string `json:"kept,omitempty"` Kept string `json:"kept,omitempty"`
// KeptMode and KeptOwner are the original's mode ("0644") and numeric owner ("0:0") as found,
// so undeclaring the file puts the original back as the machine had it (novox/hq ADR 0118).
// Absent on a record from before the host kept them; the file's mode and owner as it stands
// are used then.
KeptMode string `json:"kept_mode,omitempty"`
KeptOwner string `json:"kept_owner,omitempty"`
// Stateless is, for a service, that its unit's lifecycle was never the mesh's (novox/hq ADR // Stateless is, for a service, that its unit's lifecycle was never the mesh's (novox/hq ADR
// 0117) — kept here because removal happens once the declaration that said so is gone, and a // 0117) — kept here because removal happens once the declaration that said so is gone, and a
// service removed as if it had a state is stopped: the machine's network manager, for one. // service removed as if it had a state is stopped: the machine's network manager, for one.
Stateless bool `json:"stateless,omitempty"` Stateless bool `json:"stateless,omitempty"`
// Scope and User are, for a service in an account's own manager (novox/hq ADR 0177), which
// manager — so removal gives the unit back through the same one it was applied through.
Scope string `json:"scope,omitempty"`
User string `json:"user,omitempty"`
// Found is, for a service, the state its unit was in when this host first applied it — before // Found is, for a service, the state its unit was in when this host first applied it — before
// the mesh started, stopped, enabled or disabled anything. Removal gives that back and nothing // the mesh started, stopped, enabled or disabled anything. Removal gives that back and nothing
@@ -95,6 +105,25 @@ type Applied struct {
// file itself was — so undeclaring it gives the machine back exactly what it had. // file itself was — so undeclaring it gives the machine back exactly what it had.
Into *Into `json:"into,omitempty"` Into *Into `json:"into,omitempty"`
// Shell is, for a user, the login shell the account had before the mesh first set one, and
// the shell the mesh set last (novox/hq ADR 0176 §2, issue 228). Removal gives the found shell
// back, and only while the account still has the one the mesh set: a shell a person chose since
// is theirs. Absent when the mesh never changed the shell, and on a record written before the
// host kept it — then the shell is left exactly as it is.
Shell *LoginShell `json:"shell,omitempty"`
// Linger is, for a user, whether the account lingered before the mesh first changed it, and
// what the mesh set (novox/hq ADR 0177). Removal gives the found one back while the account
// still has the mesh's; absent when the mesh never changed it.
Linger *Lingering `json:"linger,omitempty"`
// Unpacked is, for an archive, what it put on the machine (novox/hq issue 162): the files and
// directories it unpacked, and whether the directory it was unpacked into and the parents above
// it were made by the host. Removal takes away exactly that and nothing else. Absent on a
// record written before the host kept it — then removal cannot tell the archive's files from
// anything else in the directory, and leaves it in place.
Unpacked *Unpacked `json:"unpacked,omitempty"`
// Reads is, for a container, the digest of each file it was created reading — its env-files // Reads is, for a container, the digest of each file it was created reading — its env-files
// and the files mounted into it — by path (novox/hq 04-ISSUES/103). // and the files mounted into it — by path (novox/hq 04-ISSUES/103).
// //
@@ -232,8 +261,13 @@ type FoundFirewall struct {
// retires, and returning it to adopted restores. // retires, and returning it to adopted restores.
WasActive bool `json:"was_active,omitempty"` WasActive bool `json:"was_active,omitempty"`
// DisabledByMesh is set when converging retired it, so returning to adopted enables it again // DisabledByMesh is set when converging retired it, so returning to adopted enables it again
// and nothing else ever does. // and nothing else ever does. It means exactly that (novox/hq ADR 0168): a reconcile that finds
// the firewall already inactive records RetiredBy and never this.
DisabledByMesh bool `json:"disabled_by_mesh,omitempty"` DisabledByMesh bool `json:"disabled_by_mesh,omitempty"`
// RetiredBy says how the found firewall came to be inactive on a converged machine: "mesh" when
// the mesh disabled it, "found-inactive" when a reconcile found it so and nothing of the mesh's
// had done it. Empty while it is in force or the machine is adopted.
RetiredBy string `json:"retired_by,omitempty"`
// Forward is each family's forward policy as it was before the mesh disabled the firewall, // Forward is each family's forward policy as it was before the mesh disabled the firewall,
// by the tool that sets it — recorded before, so a retirement retried puts back what the // by the tool that sets it — recorded before, so a retirement retried puts back what the
// machine had. // machine had.
@@ -592,6 +626,46 @@ type FoundUnit struct {
Boot string `json:"boot,omitempty"` Boot string `json:"boot,omitempty"`
} }
// LoginShell is what the host knows about an account's login shell, to give it back.
type LoginShell struct {
// Found is the shell the account had when the mesh first changed it. Never overwritten by a
// later change: what is given back is what was there before the mesh, not the mesh's own
// previous choice. Empty for an account the mesh created, which had no shell before it.
Found string `json:"found,omitempty"`
// Set is the shell the mesh set last — what removal compares the account against, since the
// declaration that said so is gone by then.
Set string `json:"set"`
// Created is an account the mesh made. Kept only so removal can say why there is nothing to
// give back; the account itself is never deleted.
Created bool `json:"created,omitempty"`
}
// Lingering is what the host knows about whether an account's manager runs with nobody logged in,
// to give it back (novox/hq ADR 0177).
type Lingering struct {
// Found is whether it lingered when the mesh first changed it. Never overwritten by a later
// change, as LoginShell.Found is not.
Found bool `json:"found"`
// Set is what the mesh set last.
Set bool `json:"set"`
}
// Unpacked is what an archive put under its directory, so undeclaring it takes away exactly that
// (novox/hq issue 162).
type Unpacked struct {
// Files are the files the archive placed, relative to its directory, slash-separated.
Files []string `json:"files"`
// Dirs are the directories inside it the host made for the archive — never one that was
// there before, so removal never takes a directory somebody else made, even an empty one.
Dirs []string `json:"dirs,omitempty"`
// Made is that the host made the directory itself: it was not there before the archive. One
// that was there before is never removed, empty or not.
Made bool `json:"made,omitempty"`
// Parents are the directories above it the host made to reach it, deepest first; each is
// removed on the way out only once it is empty.
Parents []string `json:"parents,omitempty"`
}
// PendingFound is a unit as found by an apply of its service that has not yet been recorded, and // PendingFound is a unit as found by an apply of its service that has not yet been recorded, and
// who asked for that apply — so only a declaration from the same origin can say it is gone. // who asked for that apply — so only a declaration from the same origin can say it is gone.
type PendingFound struct { type PendingFound struct {
+5
View File
@@ -46,6 +46,11 @@ func (a alpine) PackageInstalled(ctx context.Context, run Runner, name string) (
return strings.TrimSpace(out) != "", nil return strings.TrimSpace(out) != "", nil
} }
func (alpine) RemovePackage(ctx context.Context, run Runner, name string) error {
_, err := run(ctx, "apk", "del", name)
return err
}
func (alpine) InstallPackage(ctx context.Context, run Runner, name string) error { func (alpine) InstallPackage(ctx context.Context, run Runner, name string) error {
_, err := run(ctx, "apk", "add", "--no-cache", name) _, err := run(ctx, "apk", "add", "--no-cache", name)
return err return err
+4
View File
@@ -65,6 +65,10 @@ func (a android) InstallPackage(context.Context, Runner, string) error {
return fmt.Errorf("%w: package", ErrUnsupported) return fmt.Errorf("%w: package", ErrUnsupported)
} }
func (a android) RemovePackage(context.Context, Runner, string) error {
return fmt.Errorf("%w: package", ErrUnsupported)
}
func (a android) ServiceState(context.Context, Runner, string) (string, error) { func (a android) ServiceState(context.Context, Runner, string) (string, error) {
return "", fmt.Errorf("%w: service (init is not reachable without root)", ErrUnsupported) return "", fmt.Errorf("%w: service (init is not reachable without root)", ErrUnsupported)
} }
+199 -26
View File
@@ -2,8 +2,12 @@ package system
import ( import (
"context" "context"
"errors"
"fmt" "fmt"
"os"
"path/filepath"
"strings" "strings"
"time"
"github.com/novox/mesh-host/internal/declaration" "github.com/novox/mesh-host/internal/declaration"
) )
@@ -39,6 +43,14 @@ func (a arch) PackageInstalled(ctx context.Context, run Runner, name string) (bo
return true, nil return true, nil
} }
// RemovePackage removes one package and nothing it depends on: `-R`, not `-Rs`, because what else
// relied on a dependency is not this declaration's to know. pacman keeps a configuration file the
// operator changed as `.pacsave`, which is what "never flushed" comes to once the front end is gone.
func (arch) RemovePackage(ctx context.Context, run Runner, name string) error {
_, err := run(ctx, "pacman", "-R", "--noconfirm", name)
return err
}
func (arch) InstallPackage(ctx context.Context, run Runner, name string) error { func (arch) InstallPackage(ctx context.Context, run Runner, name string) error {
out, err := run(ctx, "pacman", "-S", "--noconfirm", "--needed", name) out, err := run(ctx, "pacman", "-S", "--noconfirm", "--needed", name)
if err == nil { if err == nil {
@@ -48,42 +60,119 @@ func (arch) InstallPackage(ctx context.Context, run Runner, name string) error {
// **The package manager's own words, and a name for the case that looks like a bug in the // **The package manager's own words, and a name for the case that looks like a bug in the
// declaration and is not.** A stale index asks the mirrors for a version they have already // declaration and is not.** A stale index asks the mirrors for a version they have already
// superseded and gets a 404 from every one of them — so the package exists, the declaration is // superseded and gets a 404 from every one of them — so the package exists, the declaration is
// correct, and the machine's idea of what exists is old (novox/hq 04-ISSUES/002). // correct, and the machine's idea of what exists is old (novox/hq 04-ISSUES/002). A keyring as
// old as the index fails one step later, on the signature of whatever a mirror still had.
//
// **Read from everything pacman said.** Its errors go to stderr, which the runner folds into
// the error rather than the output; this classifier read the output alone and so never saw a
// single "failed retrieving file", and the control node reported a ten-week-old database as
// a mirror outage with a wall of 404s (novox/hq 04-ISSUES/205).
// //
// **It is not fixed by syncing here.** `pacman -Sy <pkg>` installs a package built against // **It is not fixed by syncing here.** `pacman -Sy <pkg>` installs a package built against
// libraries this machine does not have: a partial upgrade, which Arch does not support and // libraries this machine does not have: a partial upgrade, which Arch does not support and
// which breaks the machine in a way that surfaces much later as something unrelated. The // which breaks the machine in a way that surfaces much later as something unrelated. The
// remedy is a full upgrade, and it is a decision about the whole machine rather than // remedy is a full upgrade, and it is a decision about the whole machine rather than
// something to do silently in the middle of applying one resource. // something to do silently in the middle of applying one resource. Whose decision, and on
// // what schedule, is issue 205's question; until it is answered the host says what it sees.
// So this says which of the two it is looking at. A declaration that is wrong and a machine said := strings.TrimSpace(out + "\n" + err.Error())
// that is out of date fail identically otherwise, and they are fixed in completely different switch classifyInstallFailure(said) {
// places. case installStale:
if staleIndex(out) {
return fmt.Errorf( return fmt.Errorf(
"%s could not be fetched from any mirror, which is what a stale package index looks "+ "%s could not be fetched from any mirror, which is what a stale package index looks like: "+
"like: this machine is asking for a version the mirrors have replaced. The "+ "the package database on this machine is %s and the mirrors no longer serve what it "+
"package and the declaration are probably both fine. It is fixed by upgrading "+ "names. It is fixed by upgrading the machine — a full upgrade (`pacman -Syu`) by its "+
"the machine, not by this host syncing one package — that would be a partial "+ "operator — before the mesh can install %s. The package and the declaration are probably both fine; the "+
"upgrade, which this distribution does not support.\n\n%s", "host does not sync one package by itself, because on this distribution that is a "+
name, strings.TrimSpace(out)) "partial upgrade (novox/hq 04-ISSUES/205).\n\n%s",
name, syncDatabaseAge(), name, said)
case installMirrors:
return fmt.Errorf(
"no mirror could be reached to fetch %s, and the package database on this machine is "+
"%s: this reads as the mirrors or the network, not as this machine being out of "+
"date — try again when they answer.\n\n%s",
name, syncDatabaseAge(), said)
} }
return fmt.Errorf("%w\n\n%s", err, strings.TrimSpace(out)) return fmt.Errorf("%w\n\n%s", err, strings.TrimSpace(out))
} }
// staleIndex reports whether a failed install looks like the machine's view being old rather than // How a failed install is read, from what the package manager said.
// the package being wrong. type installFailure int
//
// By what the package manager said, because there is nothing else to go on: the exit code is the const (
// same for both. installOther installFailure = iota
func staleIndex(out string) bool { // installStale: the machine's package database or keyring is older than what the mirrors
said := strings.ToLower(out) // serve — every mirror 404s the file the database names, or a package that did arrive fails
if !strings.Contains(said, "failed retrieving file") && !strings.Contains(said, "404") { // its signature against a keyring that never saw the key.
return false installStale
// installMirrors: no mirror could be reached at all, and nothing says the database is old.
installMirrors
)
// classifyInstallFailure reads pacman's words, because there is nothing else to go on: the exit
// code is the same for every one of these.
func classifyInstallFailure(said string) installFailure {
lower := strings.ToLower(said)
gone := strings.Count(lower, "returned error: 404")
fetching := strings.Contains(lower, "failed retrieving file")
badSignature := strings.Contains(lower, "invalid or corrupted package (pgp signature)") ||
strings.Contains(lower, "signature from") && strings.Contains(lower, "is invalid") ||
strings.Contains(lower, "is unknown trust") ||
strings.Contains(lower, "could not be looked up remotely")
switch {
case badSignature:
return installStale
case fetching && gone > 0:
// Every mirror, not one: a single mirror failing is an ordinary transient thing and
// retrying is the answer. pacman walks its whole mirror list before giving up, so more
// than one 404 among the lines is the index being old rather than one host being wrong.
if gone > 1 || !strings.Contains(lower, "could not resolve host") &&
!strings.Contains(lower, "connection timed out") && !strings.Contains(lower, "failed to connect") {
return installStale
} }
// Every mirror, not one. A single mirror failing is an ordinary transient thing and retrying return installMirrors
// is the answer; every one of them saying the file is gone is the index being old. case fetching:
return strings.Contains(said, "error") || strings.Count(said, "404") > 1 return installMirrors
}
return installOther
}
// staleIndex is the yes-or-no form older callers and tests use.
func staleIndex(out string) bool { return classifyInstallFailure(out) == installStale }
// syncDatabaseAge says how old this machine's package database is, in words a person acts on:
// the newest of pacman's sync databases, dated, and how long ago that was. Said beside a failed
// install so a ten-week-old database is told apart from a mirror outage by reading one line.
//
// A variable so a test can say what the machine's database looks like without having one.
var syncDatabaseAge = func() string {
entries, err := filepath.Glob("/var/lib/pacman/sync/*.db")
if err != nil || len(entries) == 0 {
return "of unknown age (no sync database found under /var/lib/pacman/sync)"
}
var newest time.Time
for _, e := range entries {
info, err := os.Stat(e)
if err == nil && info.ModTime().After(newest) {
newest = info.ModTime()
}
}
if newest.IsZero() {
return "of unknown age"
}
return describeAge(newest, time.Now())
}
// describeAge is "from 2026-07-24, 10 weeks old" — the date for the record, the span for the eye.
func describeAge(when, now time.Time) string {
days := int(now.Sub(when).Hours() / 24)
span := fmt.Sprintf("%d days old", days)
switch {
case days < 1:
span = "less than a day old"
case days >= 14:
span = fmt.Sprintf("%d weeks old", days/7)
}
return fmt.Sprintf("from %s, %s", when.Format("2006-01-02"), span)
} }
// ServiceState reads what systemd says about a unit. // ServiceState reads what systemd says about a unit.
@@ -98,7 +187,7 @@ func staleIndex(out string) bool {
// so LoadState is what is read — and it is the thing an interface spanning systemd and OpenRC // so LoadState is what is read — and it is the thing an interface spanning systemd and OpenRC
// would have had to drop. // would have had to drop.
func (arch) ServiceState(ctx context.Context, run Runner, unit string) (string, error) { func (arch) ServiceState(ctx context.Context, run Runner, unit string) (string, error) {
out, _ := run(ctx, "systemctl", "show", unit, out, asked := run(ctx, "systemctl", "show", unit,
"--property=LoadState", "--property=ActiveState", "--property=Type", "--property=LoadState", "--property=ActiveState", "--property=Type",
"--property=RemainAfterExit", "--property=ExecMainStatus") "--property=RemainAfterExit", "--property=ExecMainStatus")
@@ -124,6 +213,11 @@ func (arch) ServiceState(ctx context.Context, run Runner, unit string) (string,
switch load { switch load {
case "": case "":
// With why, where it said why: an account's manager that could not be reached says so
// here, and nowhere else (novox/hq ADR 0177).
if asked != nil {
return "", fmt.Errorf("the service manager said nothing about %s: %w", unit, asked)
}
return "", fmt.Errorf("the service manager said nothing about %s", unit) return "", fmt.Errorf("the service manager said nothing about %s", unit)
case "not-found": case "not-found":
return "", fmt.Errorf( return "", fmt.Errorf(
@@ -276,3 +370,82 @@ func (arch) ReloadService(ctx context.Context, run Runner, unit string) error {
_, err := run(ctx, "systemctl", "reload", unit) _, err := run(ctx, "systemctl", "reload", unit)
return err return err
} }
// An account's own service manager (novox/hq ADR 0177).
//
// systemd runs one manager per account beside the machine's, as user@<uid>.service, from the
// account's first login until its last logout — or for as long as the machine runs, when the
// account lingers. A user-scoped unit lives in that manager, so it can be acted on only while the
// manager runs, and lingering is what makes it run with nobody logged in.
// lingerDir is where systemd-logind keeps which accounts linger: one empty file per account. A
// variable so a test can give it a directory of its own; LingerIn is how.
var lingerDir = "/var/lib/systemd/linger"
// LingerIn points where the machine keeps lingering at a directory a test owns, until the returned
// function puts it back. Nothing outside a test calls it.
func LingerIn(dir string) (restore func()) {
was := lingerDir
lingerDir = dir
return func() { lingerDir = was }
}
// Lingering is whether an account's manager is kept running with nobody logged in. Read from
// logind's own record rather than from `loginctl show-user`, which answers only for an account
// that is logged in or already lingers — absence there is an error, and absence here is the answer.
func (arch) Lingering(_ context.Context, _ Runner, name string) (bool, error) {
_, err := os.Stat(filepath.Join(lingerDir, name))
if err == nil {
return true, nil
}
if os.IsNotExist(err) {
return false, nil
}
return false, fmt.Errorf("whether %q lingers could not be read: %w", name, err)
}
// SetLingering has logind keep an account's manager running with nobody logged in, or stop doing
// so. Disabling it stops the manager of an account that is not logged in, and every unit in it.
func (arch) SetLingering(ctx context.Context, run Runner, name string, on bool) error {
verb := "enable-linger"
if !on {
verb = "disable-linger"
}
if _, err := run(ctx, "loginctl", verb, name); err != nil {
return fmt.Errorf("cannot %s for %q: %w", verb, name, err)
}
return nil
}
// UserManagerRunning is whether an account's own manager runs now, asked of the machine's manager
// — never of the account's.
//
// **Asking the account's manager would start it.** `systemctl --user --machine=<account>@` reaches
// it through `systemd-run --machine=<account>@.host -p PAMName=login systemd-stdio-bridge`: a login,
// which starts the account's manager if none runs, for as long as that one command takes. The host
// would then act on a manager that ends a moment later and take the account's whole session with it
// — a unit started into it stops again, and the next apply starts it again. So whether there is a
// manager is read from user@<uid>.service in the machine's own, which a query does not start.
//
// exists is false for an account the machine does not have.
func (arch) UserManagerRunning(ctx context.Context, run Runner, account string) (running, exists bool, err error) {
login, exists, err := LookUpUser(ctx, run, account)
if err != nil || !exists {
return false, exists, err
}
if login.UID == "" {
return false, true, fmt.Errorf("the user database gave %q no number", account)
}
// is-active exits non-zero for every answer but "active", so the words are the answer and the
// exit is not; no words at all is a manager that did not answer.
out, err := run(ctx, "systemctl", "is-active", "user@"+login.UID+".service")
state := strings.TrimSpace(out)
if state == "" {
if err == nil {
err = errors.New("no answer")
}
return false, true, fmt.Errorf("the service manager did not say whether %q's own manager runs: %w",
account, err)
}
return state == "active", true, nil
}
+115
View File
@@ -0,0 +1,115 @@
package system
import (
"context"
"errors"
"strings"
"testing"
"time"
)
// pacman's own words from the control node on 2026-10-02 (novox/hq 04-ISSUES/205): every mirror
// 404s the versioned file a ten-week-old database names, and the one copy that arrives fails its
// signature. Errors are pacman's stderr, which the runner folds into the error, not the output.
const staleStderr = `error: failed retrieving file 'nodejs-26.5.0-1-x86_64.pkg.tar.zst' from mirror.hetzner.com : The requested URL returned error: 404
error: failed retrieving file 'nodejs-26.5.0-1-x86_64.pkg.tar.zst' from mirror.rackspace.com : The requested URL returned error: 404
error: failed retrieving file 'nodejs-26.5.0-1-x86_64.pkg.tar.zst' from arch.lucassymons.net : Could not resolve host: arch.lucassymons.net
warning: fatal error from arch.lucassymons.net, skipping for the remainder of this transaction
error: failed retrieving file 'nodejs-26.5.0-1-x86_64.pkg.tar.zst' from mirrors.cqu.edu.cn : Connection timed out after 10002 milliseconds
error: nodejs: signature from "Bert Peters (packager key) <bertptrs@archlinux.org>" is invalid
error: failed to commit transaction (invalid or corrupted package (PGP signature))`
const staleStdout = `resolving dependencies...
looking for conflicting packages...
Packages (4) ada-3.4.4-1 c-ares-1.34.8-1 simdjson-1:4.6.4-1 nodejs-26.5.0-1
:: Retrieving packages...
nodejs-26.5.0-1-x86_64 downloading...
checking keyring...
checking package integrity...
:: File /var/cache/pacman/pkg/nodejs-26.5.0-1-x86_64.pkg.tar.zst is corrupted (invalid or corrupted package (PGP signature)).
Errors occurred, no packages were upgraded.`
// A runner that behaves as ExecRunner does on failure: stdout as the output, stderr in the error.
func pacmanFailing(stdout, stderr string) Runner {
return func(_ context.Context, name string, args ...string) (string, error) {
return stdout, errors.New(name + " exited 1: " + stderr)
}
}
func TestAStaleDatabaseIsSaidAsOneWithItsAgeAndTheRemedy(t *testing.T) {
was := syncDatabaseAge
defer func() { syncDatabaseAge = was }()
syncDatabaseAge = func() string {
return describeAge(time.Date(2026, 7, 24, 16, 56, 0, 0, time.UTC), time.Date(2026, 10, 3, 0, 0, 0, 0, time.UTC))
}
err := arch{}.InstallPackage(context.Background(), pacmanFailing(staleStdout, staleStderr), "nodejs")
if err == nil {
t.Fatal("a failed install must fail")
}
for _, want := range []string{
"the package database on this machine is from 2026-07-24, 10 weeks old",
"stale package index",
"upgrading the machine — a full upgrade (`pacman -Syu`) by its operator — before the mesh can install nodejs",
"partial upgrade",
"returned error: 404", // pacman's own words follow
} {
if !strings.Contains(err.Error(), want) {
t.Errorf("the error does not say %q:\n%s", want, err)
}
}
if strings.Contains(err.Error(), "mirrors or the network") {
t.Errorf("a stale database must not be read as a mirror outage:\n%s", err)
}
}
// The same 404s read from stderr alone — the half this classifier used to be blind to.
func TestTheClassifierReadsWhatPacmanWroteToStderr(t *testing.T) {
if classifyInstallFailure(staleStderr) != installStale {
t.Fatal("every mirror 404ing the named file is a stale database")
}
if classifyInstallFailure(staleStdout) != installStale {
t.Fatal("a corrupted-signature line alone is a stale keyring")
}
if classifyInstallFailure("") != installOther {
t.Fatal("nothing said is nothing classified")
}
}
func TestAnUnreachableMirrorWithAFreshDatabaseIsAMirrorProblem(t *testing.T) {
was := syncDatabaseAge
defer func() { syncDatabaseAge = was }()
syncDatabaseAge = func() string { return "from 2026-10-02, less than a day old" }
outage := `error: failed retrieving file 'core.db' from mirror.hetzner.com : Could not resolve host: mirror.hetzner.com
error: failed retrieving file 'core.db' from mirror.rackspace.com : Connection timed out after 10001 milliseconds
error: failed to synchronize all databases (failed to retrieve some files)`
if classifyInstallFailure(outage) != installMirrors {
t.Fatal("no mirror answering, no 404, no signature fault: the mirrors, not the machine")
}
err := arch{}.InstallPackage(context.Background(), pacmanFailing("", outage), "nodejs")
if err == nil || !strings.Contains(err.Error(), "mirrors or the network") || !strings.Contains(err.Error(), "less than a day old") {
t.Errorf("a mirror outage is said as one, with the database's age beside it:\n%v", err)
}
}
func TestAFailureThatIsNeitherKeepsPacmansWords(t *testing.T) {
err := arch{}.InstallPackage(context.Background(), pacmanFailing("", "error: target not found: nodejsx"), "nodejsx")
if err == nil || !strings.Contains(err.Error(), "target not found") || strings.Contains(err.Error(), "package database on this machine") {
t.Errorf("an unknown package is pacman's own error, not a stale database:\n%v", err)
}
}
func TestDescribeAge(t *testing.T) {
now := time.Date(2026, 10, 3, 0, 0, 0, 0, time.UTC)
for when, want := range map[time.Time]string{
now.Add(-2 * time.Hour): "less than a day old",
now.Add(-5 * 24 * time.Hour): "5 days old",
now.Add(-71 * 24 * time.Hour): "10 weeks old",
} {
if got := describeAge(when, now); !strings.HasSuffix(got, want) {
t.Errorf("%s: got %q, want suffix %q", when, got, want)
}
}
}
+40
View File
@@ -0,0 +1,40 @@
package system
import (
"context"
"errors"
"os/exec"
"testing"
)
// A user that does not exist yet is "absent", in the words the host's own runner uses
// ("getent exited 2"), not only Go's ("exit status 2"). Matched as text, the runner's wording read as
// a user database that did not answer, and the controller's account was never created.
func TestAMissingUserIsAbsentInTheRunnersWords(t *testing.T) {
for _, words := range []string{"getent exited 2: ", "exit status 2"} {
run := func(context.Context, string, ...string) (string, error) { return "", errors.New(words) }
_, found, err := LookUpUser(context.Background(), run, "nobody-here")
if err != nil || found {
t.Errorf("%q: found %v, err %v; want absent", words, found, err)
}
}
run := func(context.Context, string, ...string) (string, error) { return "", errors.New("getent exited 1: ") }
if _, _, err := LookUpUser(context.Background(), run, "x"); err == nil {
t.Error("a database that failed read as an answer")
}
}
// The real command, through a real exit: getent's code for a key not found.
func TestAMissingUserIsAbsentFromTheRealGetent(t *testing.T) {
if _, err := exec.LookPath("getent"); err != nil {
t.Skip("no getent here")
}
run := func(ctx context.Context, name string, args ...string) (string, error) {
out, err := exec.CommandContext(ctx, name, args...).Output()
return string(out), err
}
_, found, err := LookUpUser(context.Background(), run, "mesh-no-such-user-0b1f")
if err != nil || found {
t.Errorf("found %v, err %v; want absent", found, err)
}
}
+94 -3
View File
@@ -19,6 +19,11 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"os"
"os/exec"
"path/filepath"
"regexp"
"strconv"
"strings" "strings"
"github.com/novox/mesh-host/internal/declaration" "github.com/novox/mesh-host/internal/declaration"
@@ -51,6 +56,9 @@ type System interface {
PackageInstalled(ctx context.Context, run Runner, name string) (bool, error) PackageInstalled(ctx context.Context, run Runner, name string) (bool, error)
InstallPackage(ctx context.Context, run Runner, name string) error InstallPackage(ctx context.Context, run Runner, name string) error
// RemovePackage uninstalls one package, leaving its dependencies and anything the operator
// changed in its configuration where the package manager leaves them (novox/hq ADR 0180).
RemovePackage(ctx context.Context, run Runner, name string) error
// ServiceState is "running" or "stopped". A unit that does not exist is an error, never // ServiceState is "running" or "stopped". A unit that does not exist is an error, never
// "stopped" — reporting absence as satisfaction is the fault this host exists to prevent. // "stopped" — reporting absence as satisfaction is the fault this host exists to prevent.
@@ -76,6 +84,9 @@ type System interface {
type Login struct { type Login struct {
Home string Home string
Shell string Shell string
// UID is the account's number, as the user database gives it: what names the account's own
// service manager to the machine's (user@<uid>.service, novox/hq ADR 0177).
UID string
} }
// LookUpUser reads a login from the user database. // LookUpUser reads a login from the user database.
@@ -92,8 +103,10 @@ func LookUpUser(ctx context.Context, run Runner, name string) (Login, bool, erro
out, err := run(ctx, "getent", "passwd", name) out, err := run(ctx, "getent", "passwd", name)
if err != nil { if err != nil {
// getent's own convention: 2 means the key was not found, which is the only failure that // getent's own convention: 2 means the key was not found, which is the only failure that
// means "no such user". // means "no such user". Asked of the exit code: matched as text, it looked for Go's wording
if strings.Contains(err.Error(), "exit status 2") { // ("exit status 2") while the host's runner says "getent exited 2", so a user that did not
// exist yet read as a database that did not answer, and no account was ever created.
if code, ok := ExitCode(err); ok && code == 2 {
return Login{}, false, nil return Login{}, false, nil
} }
return Login{}, false, fmt.Errorf( return Login{}, false, fmt.Errorf(
@@ -105,7 +118,7 @@ func LookUpUser(ctx context.Context, run Runner, name string) (Login, bool, erro
return Login{}, false, fmt.Errorf("the user database gave %q for %q, which is not a passwd entry", return Login{}, false, fmt.Errorf("the user database gave %q for %q, which is not a passwd entry",
strings.TrimSpace(out), name) strings.TrimSpace(out), name)
} }
return Login{Home: fields[5], Shell: fields[6]}, true, nil return Login{Home: fields[5], Shell: fields[6], UID: fields[2]}, true, nil
} }
// GroupsOf is every group a login is in. // GroupsOf is every group a login is in.
@@ -117,6 +130,60 @@ func GroupsOf(ctx context.Context, run Runner, name string) ([]string, error) {
return strings.Fields(out), nil return strings.Fields(out), nil
} }
// shells is where the machine lists the shells a login may have (shells(5)). A variable so a test
// can point it at a list of its own; ShellsIn is how.
var shells = "/etc/shells"
// ShellsIn points where the machine's shells are listed at a file a test owns, until the returned
// function puts it back. Nothing outside a test calls it.
func ShellsIn(list string) (restore func()) {
was := shells
shells = list
return func() { shells = was }
}
// UsableShell says why a path cannot be an account's login shell, or nil when it can.
//
// **Asked before a shell is set, because nothing after it would say.** `usermod --shell` only
// warns about a shell that is missing or not executable, and succeeds; the host's read-back
// compares the user database's string, which then matches. So an account could be pointed at a
// shell that is not there, and console, ssh and display-manager logins all fail — after a
// failed package install, say, which does not stop the resources after it (novox/hq issue 228).
//
// Listed among the machine's shells as well as executable, because that list is what login
// services check: an unlisted shell is one ssh and the display manager may refuse.
//
// **Except a shell that refuses a login.** nologin and false are how a service's account says it
// is not a login at all, and no distribution lists them — the controller's own account has one.
// Requiring them listed would refuse every service account; they are still required to exist.
func UsableShell(path string) error {
if !filepath.IsAbs(path) {
return fmt.Errorf("the shell %q is not an absolute path", path)
}
info, err := os.Stat(path)
if err != nil {
return fmt.Errorf("the shell %s is not on this machine: %w", path, err)
}
if !info.Mode().IsRegular() || info.Mode().Perm()&0o111 == 0 {
return fmt.Errorf("the shell %s is not an executable file", path)
}
if base := filepath.Base(path); base == "nologin" || base == "false" {
return nil
}
raw, err := os.ReadFile(shells)
if err != nil {
// Unreadable is not "not listed": the two are told apart, as the user database's are.
return fmt.Errorf("the machine's shells (%s) could not be read, so %s cannot be checked: %w",
shells, path, err)
}
for _, line := range strings.Split(string(raw), "\n") {
if line = strings.TrimSpace(line); line == path {
return nil
}
}
return fmt.Errorf("the shell %s is not listed in %s, so logins may refuse it", path, shells)
}
// Supports reports whether this host can apply a shape. // Supports reports whether this host can apply a shape.
func Supports(s System, t declaration.Type) bool { func Supports(s System, t declaration.Type) bool {
for _, shape := range s.Shapes() { for _, shape := range s.Shapes() {
@@ -218,3 +285,27 @@ func For(name string) (System, error) {
func All() []System { func All() []System {
return []System{arch{}, alpine{}, android{}} return []System{arch{}, alpine{}, android{}}
} }
// ExitCode is the code a command exited with, when err says one: from the exit itself where the
// runner kept it, else from the words either runner shape uses ("exit status N", "<cmd> exited N").
func ExitCode(err error) (int, bool) {
if err == nil {
return 0, false
}
var exit *exec.ExitError
if errors.As(err, &exit) {
return exit.ExitCode(), true
}
if m := exitWords.FindStringSubmatch(err.Error()); len(m) == 3 {
for _, g := range m[1:] {
if g != "" {
if n, convErr := strconv.Atoi(g); convErr == nil {
return n, true
}
}
}
}
return 0, false
}
var exitWords = regexp.MustCompile(`exit status (\d+)|exited (\d+)`)
+75
View File
@@ -0,0 +1,75 @@
package system
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
// novox/hq ADR 0177: whether an account's own manager runs is asked of the machine's manager, by
// the account's number — never of the account's, which the question would start.
func TestAnAccountsManagerIsAskedOfTheMachinesManager(t *testing.T) {
var asked []string
up := false
run := func(_ context.Context, name string, args ...string) (string, error) {
line := name + " " + strings.Join(args, " ")
asked = append(asked, line)
switch {
case line == "getent passwd ops":
return "ops:x:1001:1001::/home/ops:/bin/bash\n", nil
case name == "getent":
return "", errors.New("getent exited 2: ")
case line == "systemctl is-active user@1001.service":
if up {
return "active\n", nil
}
return "inactive\n", errors.New("systemctl exited 3: ")
}
t.Fatalf("asked %q", line)
return "", nil
}
a := arch{}
for _, want := range []bool{false, true} {
up = want
running, exists, err := a.UserManagerRunning(context.Background(), run, "ops")
if err != nil || !exists || running != want {
t.Fatalf("up %v: running %v exists %v err %v", want, running, exists, err)
}
}
if _, exists, err := a.UserManagerRunning(context.Background(), run, "nobody-here"); err != nil || exists {
t.Fatalf("an account the machine does not have: exists %v err %v", exists, err)
}
for _, c := range asked {
if strings.Contains(c, "--user") || strings.Contains(c, "--machine") {
t.Fatalf("the account's own manager was asked: %s", c)
}
}
}
func TestLingeringIsReadFromLogindsRecordAndSetThroughLoginctl(t *testing.T) {
dir := t.TempDir()
defer LingerIn(dir)()
a := arch{}
if on, err := a.Lingering(context.Background(), nil, "ops"); err != nil || on {
t.Fatalf("no record read as lingering: %v %v", on, err)
}
if err := os.WriteFile(filepath.Join(dir, "ops"), nil, 0o644); err != nil {
t.Fatal(err)
}
if on, err := a.Lingering(context.Background(), nil, "ops"); err != nil || !on {
t.Fatalf("logind's record not read as lingering: %v %v", on, err)
}
var asked []string
run := func(_ context.Context, name string, args ...string) (string, error) {
asked = append(asked, name+" "+strings.Join(args, " "))
return "", nil
}
_ = a.SetLingering(context.Background(), run, "ops", true)
_ = a.SetLingering(context.Background(), run, "ops", false)
if strings.Join(asked, "; ") != "loginctl enable-linger ops; loginctl disable-linger ops" {
t.Fatalf("%v", asked)
}
}