19e5dd83ea1c2454027501909b8b85ae8c48df9c
26
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
19e5dd83ea |
apply: a run-once container is a step the host runs to completion (ADR 0052)
A module can declare state but not a step that runs at first boot. This adds `run-once: true` to the container shape: the host runs it in the foreground, requires it to exit 0, and records that it did — as the digest of the declaration, so a re-apply does not re-run it unless the declaration changed. Because the declaration is applied in order and a failed run-once step gates the apply the way a failed action does, whatever is declared after the step starts only once it has completed. That is how "before the broker starts" is enforced, with no dependency graph the host must resolve (ADR 0005): the step is declared first, and the container that needs it is never reached until it is done. No new host shape and no arbitrary host command — a run-once container is strictly less powerful than an action. Validation refuses run-once with restart-on (contradictory lifecycles). Six unit tests; go test ./... green. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
f06eea5fa3 |
declaration: an access is mounted, and the host owns nothing about it
The tenth shape (novox/hq ADR 0051). Shared, pre-existing data — a media library, a download spool several modules use — is the operator's, not the mesh's. A `directory` resource is the host's own: it creates it, chowns it, sets its mode and removes it when empty. An access is the opposite on every axis. Add the `access` type to the vocabulary. Its applier confirms the path is present and changes nothing: it does not create, chown, reconcile or set a mode. Absent is refused clearly — the operator must provide it — rather than created, because a bind mount whose source is missing is made as root by the container runtime with the wrong ownership (04-ISSUES/026). Undeclaring an access forgets the record and never touches the path, which is the data loss ADR 0030 prevents, on a directory the mesh never made. Full hosts speak it (it gates a bind mount, which needs the container runtime); the vocabulary guard test records the decision that made it the tenth shape. Unit tests cover present, absent-refused, and undeclared-left-alone. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
aa441bac19 |
A container reflects its config: restart-on for containers (04-ISSUES/009)
A container reads a mounted file once, at start; its spec (image, env, volumes) does not include a mounted file's content, so a settings change that re-renders the file left the running process holding the old value while every check passed. Give Container the restart-on field a Service already has, and recreate the container when a named resource changed this pass. Unit-tested (recreated on change, left alone otherwise) and proven in the mesh-lab: a running grafana runtime picked up a token change on the next push. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
aca9eb37ff |
An owner may be a number the machine has never heard of
A directory a module mounts into its container belongs to whoever runs inside — grafana's 472, redis's 999, www-data's 33 — and none of those has a row in the machine's passwd. Owner-by-name refused them all, which looked principled and meant every module whose container drops privileges could not own its own data. The lab showed both coats of it in one run: the store's config file was unreadable to the store, restarting forever on permission denied, and the forge could not traverse into the 0700 root-owned directory that held its files — a directory that had only become root-owned when declaring it fixed 04-ISSUES/026, because Docker used to create it 0755. A fix that tightens ownership without a way to say whose it should be moves the fault, not removes it. "uid:gid" and bare "uid" are numeric and chowned as given; a name still resolves as before, and a name with a colon is refused rather than half-read. |
||
|
|
b91342a6bd |
A machine says which ports it already holds
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a node raises it from the bundle it carries before any mesh exists, so the control plane has never heard of the store, the broker, or the control plane's own container. A module assigned afterwards is handed a port one of them holds, and finds out from a container runtime three layers down. The host already recorded which resources it carried and which the mesh sent — that distinction exists so the two never remove each other. It now also records what each one binds, and reports the carried ones. What the declaration binds, not what is open. A machine's open ports are a moving target — something a person started, a connection the kernel handed out — and assigning around those would mean a port that was free when it was asked for and taken when it was used. What a resource declares is stable, and it is the half the mesh can be responsible for. Only the carried ones are reported. What the mesh put here it already knows, and reporting it back would make the machine an authority on the mesh's own bookkeeping. |
||
|
|
8c248e3d7f |
A secret can reach a container's environment, and sit inside a config file
Two gaps found by writing the first real module's manifest rather than
by reasoning about one. Both are fields on existing shapes, so the
vocabulary is still nine.
**env-file on a container.** A declaration reaches a node over the
broker and `env` is plain text in it, so a password there is a password
the broker sees — the transitive trust refused everywhere else. A sealed
file arrives unreadable, the host writes it, the runtime reads it. It is
also simply how third-party software takes credentials: nothing shipping
in a container will read a path the mesh invented, and every one of them
reads its environment.
**secrets in a file's content.** A program wanting its token inside a
JSON document cannot be handed a file that is entirely a token, and the
mesh cannot compose the document because it discarded the value. So the
module supplies the document with `${secret:name}` in it, the mesh
delivers the value sealed, and the host is the only thing that ever
holds both.
Substitution is textual and the host learns no formats. Deliberate: a
mechanism that understood JSON would be asked to understand YAML next,
and then INI, which is how the arrangement this replaces became
something nobody could hold in their head. The module knows its own
format because it wrote the rest of the file. The sharp edge is stated
rather than left to be discovered — a value containing a quote is not
escaped for whatever surrounds it.
Refused in both directions, because both are somebody being wrong about
where a credential is: a placeholder with nothing to fill it would write
`${secret:x}` into a config file, and a secret the content never uses
means somebody believes a credential is in a file where it is not.
A file that carries one is 0600 unless the module said otherwise.
|
||
|
|
f48e06473d |
A directory holding anything the mesh did not put there is never removed
Found by asking what the conversion needs, and it is the one failure in this system that cannot be undone. Unassigning a module made its directory an orphan, and an orphan directory was deleted with everything under it — os.RemoveAll — while the report said "removed". A database's files, a mail spool, somebody's uploads. Reproduced before fixing: assign a module, let a service write into its directory, unassign the module, and the file is gone. Now a directory that still holds something is kept and said so, naming how many items are in it. What makes that safe rather than merely cautious is the removal order, which was already right. Everything the mesh puts in a directory is itself a declared resource, and orphans are removed in reverse declaration order — so what the mesh wrote is already gone by the time the directory is reached. Anything still there was put there by something else, which is the definition of data. It is the host's own line applied to the one shape where getting it wrong does not recover: it removes what it made and leaves what it merely configured. An empty directory is what it made; a full one is not, and an empty one is still removed so nothing accumulates. Files are unchanged. A declared file is the mesh's own, and losing a config file is not the failure this is about. |
||
|
|
4a43e21794 |
A network is a shape, so that it can be removed
novox/hq ADR 0029, and work breakdown 1.3. A module of several containers had no way to let them reach each other by name: a container declaration could join a network and nothing could create one. An action was the obvious alternative and is refused on removal — "an action has no footprint the host can undo", so a network made that way outlives every module that is ever unassigned, and the mesh cannot tell. A resource the mesh can create and never clean up is one it should not create. A name and nothing else. Not a driver, a subnet or a gateway: each is something a module would have to know about the machine it lands on, and a module naming a subnet collides with whatever else chose the same one. It needs no new ordering rule. Resources apply in declaration order and orphans are removed in reverse, so a network written before the containers that join it is created first and removed last — after they are gone. A runtime refusing to remove one still in use is reported rather than swallowed, because that means something undeclared is holding it. The vocabulary guard fired on the change, as designed, and now names the record instead of a number: nine shapes, with the argument beside the count. Creation reads back rather than trusting an exit status (ADR 0018): a runtime that reports success and made nothing leaves every container that joins it failing to start, one step from the cause. |
||
|
|
af9d316258 |
Resources are applied in the order they were declared, and now something says so
Half of novox/hq work breakdown 1.3, and it needed no change: the apply loop walks d.Resources and sorts nothing, so a module that needs one thing before another says so by writing it first. Asserted because it is the kind of property a later change breaks silently. Sorting the resources for any good reason at all — by type, by identity, for a tidier report — would still pass every other test in this package. It is sequence, not readiness. A container started is not a container ready, and nothing here waits: what depends on something being usable retries, which is what both example provisioners do and is the more robust answer anyway, because a dependency can restart long after everything was applied. Two mistakes worth keeping in the test's own comments. The first version stubbed the runner to always succeed, so verify passed, every action counted as already done, and nothing ran — the assertion was measuring an empty list. The second declared the actions over the link, which refuses them: only a bundle may carry an action (ADR 0005). |
||
|
|
8e12b3c9e4 |
Name the decisions these tests defend, and check the bundle at all
From auditing the decision records: of 28, only 12 were named by any test, so "which decisions are defended" could not be answered without reading everything. ADR 0017 says a test names the decision it defends — that rule was itself unenforced. Most of the gap was citation, not coverage. Drift detection was tested in several places without naming ADR 0011; the archive refusal without naming 0012; forged declarations without naming 0002. Named now, so the question is answerable by grep. The bundle was the real gap: nothing tested substrate-first-node.lock at all. It is what a machine becomes when there is no mesh to ask — the one declaration applied with nothing to verify it against — and it was edited by hand and read by nothing but a running host. Two tests now assert what it carries: exactly postgres, lavinmq and the control plane. That defends ADR 0028, which removed the object store from the substrate after it had been a member for months on the strength of "it cannot grant itself a bucket" — true, and the answer to only half the test. Nothing counted what the bundle held. Fault-injected, and the first attempt did not bite: the injection landed on a comment line, which stripComments discards. Injecting into the image field fails as it should. |
||
|
|
c3d6f240fe |
Give a container the names, rather than a resolver to ask
The commit before this said "told where to resolve names" and passed --dns, which is not what it ended up doing. This is that correction: a container is given the names themselves, written into its own hosts file by the runtime. The reason for the change is the decision the mesh already made about names — a file rather than a resolver, because it works on every runtime, needs no package and has no failure mode of its own. Passing a resolver address would have required a resolver to exist, which at that point none did. A resolver is coming, for the case a file genuinely cannot express: a service named under a machine, postgres.novox.internal, where the wildcard cannot be enumerated in advance. When it arrives it will need this field back under its own name. It is not being kept in the meantime — a field nothing fills is a field nobody can trust, and the vocabulary is asserted by a count for exactly that reason. |
||
|
|
0e2b288bb6 |
A container can be told where to resolve names
A container does not inherit the machine's names. It gets its own /etc/hosts holding its own hostname, and a runtime rewrites resolv.conf — so every internal name the mesh wrote for that machine is invisible to what the machine is running. That was hit for real, in the lab: a database client on one node could not resolve another node, on a mesh where both names were correct and present on both machines. It was worked around by resolving on the host and passing an address, which is the kind of workaround that should not be needed twice. A field on an existing shape, not a ninth shape — the vocabulary is still the eight the count asserts. Per container rather than by editing the machine's resolver configuration: that file belongs to something else on most machines, and a host that edited it would be fighting whatever owns it on every boot — the fault this host exists to avoid, in the place it would be hardest to see. A container told nothing is run exactly as before. Most containers should resolve whatever the machine resolves, and passing an empty flag would be a change of behaviour dressed up as a default. |
||
|
|
08e91065b4 |
A failed action stops what follows; nothing else does
The previous commit continued past every failure, and the lab found the cost immediately: the bootstrap's store-readiness gate failed, the apply carried on and started the broker and control plane against a machine that was not ready, and the database still initialising was shut down. An action is the only shape whose purpose is to make something true before the next thing needs it — which is why it is the only one with a verify. The bootstrap is a row of them. Everything else is independent state, and stopping there is what made one broken module hold a whole machine hostage. The report says which happened: "these things failed" and "these things failed and the rest was never tried" are different machines. |
||
|
|
aec37bf89e |
Attempt every resource, and report every failure
Found in the lab while proving something else. A machine assigned a module declaring a package that does not exist applied NOTHING on every later push, for ever — the broker's queues were empty, so the declaration had been delivered and read; the machine stopped at the first failing resource and never reached the rest. A machine with one bad module and nine good ones ran none of the nine, and the mesh reported "failed" without saying the rest were never attempted. Nothing that re-pushes to machines that are behind could recover it either: it would retry a permanent failure for ever and make no progress on anything else. And which nine a broken module blocks is an accident of resolution order. The behaviour had a test asserting it, citing ADR 0010. That record does not decide this — it argues about pipelines against reconcilers, and says nothing about whether one resource failing should stop the next being attempted. The citation was doing more work than the record supports. So: everything is attempted, every failure is reported, and the first line says how many. The case for stopping was that a later resource may depend on an earlier one. It still may — and it then fails its own check and is reported, which is more information than skipping it. This host reads back after every write precisely so that is caught rather than assumed. Unchanged: a declaration that cannot be PARSED is still refused whole. That is a different thing — "this machine could not do it" against "this was never a declaration" — and they are fixed in different places. Recorded as novox/hq 04-ISSUES/011 with the evidence. |
||
|
|
c57087d75d |
A user, bytes, and an archive — because most of what people install is
not a service A shell, a terminal, a chat client, a desktop are a package plus configuration in somebody's home. A mesh with no notion of a user can own /etc and nothing anybody looks at, which is most of the reason to manage a machine at all. Three shapes, and the vocabulary test asserts the count precisely because widening it widens what a compromised control plane can express: user a login, its shell and its groups archive a set of files, fetched by digest and unpacked (file) gains `bytes` for what is not text, and `owner` `user` also makes "zsh is my login shell" declared state. chsh is a command, the link may not carry one, and a shell settable only by hand is a shell the mesh cannot manage. Groups are additive and never pruned — usermod without --append REPLACES them, which would silently remove every group that makes a login able to use the machine. A machine's own groups are not the mesh's to know about. The archive is the one place this host reaches out on its own; everywhere else it holds one outbound connection and fetches nothing. So it carries the discipline the bootstrap already uses for images: pinned by digest, and the digest checked before a single file is written. Two decisions in the unpacker worth naming: - an entry naming a path outside the archive is REFUSED, not sanitised. Rewriting it to land inside would put a file somewhere nobody asked for and report success. Found by the test: the first version quietly relocated it. - symlinks and device nodes are refused rather than skipped, or an archive that needed one arrives silently incomplete. A partial host does archives and refuses users: an archive needs a filesystem and a way to fetch; a user needs a user database it is allowed to write. |
||
|
|
a752fc514b |
A file the mesh can deliver and cannot read
Everything else in a declaration is visible to whatever carried it. The message is signed so it cannot be forged, and signing does not make it unreadable — a password in `content` is a password the broker sees, which is the transitive trust this design refuses everywhere else. So a node generates a third key at enrolment and reports the public half, exactly as it does for its identity and its overlay key. A file may arrive `sealed` instead of `content`; the host opens it with that key and writes the result. The control plane can then store a credential it cannot use, and the broker relays a blob it cannot read. A third key rather than reusing one of the two. The identity key signs and is Ed25519; the overlay key is WireGuard's and is tied to being on the private network, which a machine may not be. A key used for two purposes is one rotation away from breaking the other. Details that are not incidental: - sealed and content together is refused, so "was this the secret or the placeholder" is answerable by looking - a sealed file defaults to 0600 rather than 0644, because the consequence differs; an explicit mode still wins - a node with no sealing key refuses the file rather than skipping it. A machine that quietly omits the one resource carrying a credential looks configured and cannot connect - what is recorded is a digest of what was written, so drift on a credential is still detected without the node keeping the value, and the report that goes back over the broker carries neither The key is made at enrolment rather than on first use. One made later is one the mesh was never told about, so nothing could ever be sealed to it, and the node would look fine and receive nothing. This is why sealing was borrowed from another mesh's mistakes rather than its design: there, credentials sit encrypted in the control plane's database — which guards the database file and nothing else, since the same value is also in each node's environment file in plain text and inside every connection string composed from it. Its own tooling has to search by value rather than by name to find the copies, and says the ones inside composed URLs are usually the only copies in use. |
||
|
|
827ce481f2 |
Somebody editing a managed file is now visible instead of mysterious
Asked how the mesh would know if somebody edited their hosts file. It would not. The file was rewritten within five minutes and the outcome said "updated" -- which is exactly what the mesh changing its own mind looks like. So the change vanished, nothing anywhere said why, and the obvious thing to do is edit it again. The host now records a digest of what it wrote, which is enough to tell the two apart on the next pass: the file matches the declaration unchanged it matches what was last written updated -- the mesh changed its mind it matches neither corrected -- somebody changed it here The machine is put back either way, because holding it to what it was told is the point. What changes is that it says so. A digest rather than the content: the store is read on every reconcile and sits beside the state on disk, and keeping every managed file twice would make it grow with the size of the machine rather than with the number of resources. |
||
|
|
1bc97ed50d |
A service can be declared to reflect a file
Because a running service does not re-read its configuration. Replace the file, find the service running, do nothing -- and the machine keeps behaving as it did while every check passes, because the file is right and the service is up. That is not hypothetical. It is how a third node joining a mesh left the first two carrying a private network that no longer existed, with every part of it reporting success. Declared state rather than a command: the declaration says the running service must reflect these files, and the host works out that it does not. A command to restart would be an action, and the link may not carry one -- the host refused precisely that when I tried it, correctly, which is how this shape was arrived at rather than the other. Scoped to one apply. A change from an earlier one has already been reflected, and restarting for it every time would make a steady machine bounce its services for ever. Also: the node generates its overlay key at enrolment and reports the public half, and the store waits three minutes rather than one for the database -- sixty seconds is not enough for a cold machine running initdb, and it failed that way three times, which is the worst kind of flake because a second run always fixed it. |
||
|
|
fa48b5825e |
The bundle and the mesh stop removing each other
04-ISSUES/010. The store now records where each resource came from -- carried, or declared -- and each origin removes only its own. A declaration removes what the mesh previously declared and never what the bundle raised. State written before the field existed reads as carried, because everything a host had applied by then came from its bundle: there was no other way to tell it anything. Guessing the other way would have the first upgrade remove the substrate, which is this fault arriving through the change that fixes it. Verified on the scenario that caused it, and on the property that had to survive it: a later declaration dropping a resource still removes that resource, so removal by omission still means what it meant. Also stops swallowing a publish failure. A node that applied a declaration and could not tell the mesh looked exactly like one that had -- the mesh believing it never answered, the node believing it did, and nothing anywhere saying so. Reports are published mandatory now, so anything the broker cannot route comes back and is said out loud rather than dropped in silence. |
||
|
|
a4445f5c0a |
A machine joins the mesh it raised
The last step of the first-node path, and the bundle now carries all of it: a container runtime, the store, a database per context, their schemas, the broker with a certificate it generated itself, and the control plane running. Then the machine enrols against the mesh on its own disk. It dials the broker over TLS, refuses anything but the pinned certificate, presents the one-time secret with a public key it generated, and is told the name the mesh has for it. Its specialness lasted two commands, which is what ADR 0004 asked for. The identity is saved only after the mesh says it knows this node. A node holding an identity the mesh never recorded would believe it had joined and be believed by nobody, which is worse than not joining because nothing looks wrong. An already-enrolled machine refuses a valid token rather than quietly acquiring a second identity, and a spent token is refused by the mesh. Both checked. Containers gained a network field. The control plane must reach the store and the broker on the machine it was raised on, before there is any mesh to arrange that; the alternative was publishing ports and guessing an address that works from inside a container, which fails in a worse way. The control plane talks to the broker over loopback in plaintext, deliberately. The TLS on 5671 exists so a node crossing a network can pin a certificate, not for a hop that never leaves the machine. Verified on a sealed lab machine: eleven resources applied from bare, the control plane consuming, a token issued from inside it, and the machine enrolled -- with the recorded public key matching what the host printed, the token marked spent, and the profile stored. |
||
|
|
ee2648188d |
Repoint ADR references after HQ consolidated 65 records to 23
96 comments across the two repos named records that no longer exist. Each now points at the consolidated record that holds its reasoning -- ADR 0034 (a test defends a decision) is 0017, the eight host records are 0005, the four lab records are 0016. Worth noting for next time: these are references from outside HQ, so renumbering there is not free. It cost 38 files here. |
||
|
|
02f1fcc865 |
Three hosts: arch, alpine and android
ADR 0060, built. `make hosts` produces mesh-host-arch, mesh-host-alpine and mesh-host-android, each pinned to its system at link time. The claim that "almost all of it is shared" held up. All 36 existing apply tests pass unchanged -- the only edit was naming which system they run against, which was previously implicit. What moved into internal/system is two appliers' worth of code and the probes that go with them. Each system's differences are real and needed re-deriving rather than translating: apk reports absence by EMPTY OUTPUT and exits zero either way, where pacman exits non-zero. Reading apk's exit code the way pacman's is read reports every package as installed. That is the single most dangerous difference between the two and it is invisible until it bites. OpenRC has no LoadState, so "the service does not exist" is read from its prose rather than a field. Same distinction, different evidence -- and this is exactly what an interface spanning both would have had to drop, which is why 0060 rejected one. OpenRC has no is-enabled either. Boot state comes from the runlevel listing: "does it start at boot" becomes "does it appear in rc-update show default". Android is a partial host and that is the point. It implements file, directory and action -- the shapes needing only a filesystem and a way to run something -- and refuses the other three by name, before anything is applied. Its unreachable appliers return ErrUnsupported rather than a zero value, so "unreachable" fails loudly if it stops being true. A host also confirms it is on the machine it was built for, once, at the start. The alpine host on this Arch machine says "this machine is not Alpine" instead of failing later inside a package manager that is not there. And a host built without -X main.builtFor refuses everything, naming the hosts that exist. Two test problems found by injecting faults. One injection did not compile, so the check now reports that separately from a pass. The other passed with the behaviour removed: the missing-service assertion matched "does not exist", which the FALL-THROUGH error also contains because it echoes the raw output. It now asserts the diagnosis, which only the correct branch produces. Verified with the real binaries: android refuses a package naming what it does support; alpine on Arch refuses the machine; arch applies and is idempotent; a system-less build refuses everything. |
||
|
|
f04294c3c1 |
A service can be enabled at boot, and a container uses the runtime the machine has
Two gaps found by testing podman rather than reasoning about it.
The service shape could not say "starts at boot". It ran `systemctl start`, so
`service: docker.service, running` started docker now and it would not come
back after a reboot unless something else had enabled it. A declaration that
reports success and stops being true at the next power cut.
`boot: enabled|disabled` is now a separate field, not a fourth value of
`state`, because the two are orthogonal: a unit can be enabled and stopped (it
returns at boot) or disabled and running (started by hand, gone after one).
Absent means the host asserts nothing, so a machine whose operator enabled
something is not silently disabled by a declaration that never mentioned it.
Boot state is made true BEFORE the unit is started. When an apply fails part
way, enabled-and-stopped comes back at the next boot and running-and-disabled
does not, so the more durable half goes first.
`is-enabled` has the same trap as `is-active` had. Its exit code is non-zero
for nearly everything, and `static` is neither enabled nor disabled -- the unit
has no install section and CANNOT be enabled. Reading it as "disabled" would
have the host try, fail, and blame the wrong thing, which is the same shape as
reading a missing unit as "stopped".
The container applier no longer calls `docker` literally. Verified on this
machine against podman 6.1.0:
docker info --format '{{.ServerVersion}}' -> 29.7.2
podman info --format '{{.ServerVersion}}' -> Error: can't evaluate field
ServerVersion
podman info --format '{{.Version.Version}}' -> 6.1.0
So one probe cannot find both, and a host using docker's would report a machine
running podman as having no container runtime at all. Everything else IS
compatible -- run, rm -f, and docker's own Go template syntax for reading state
and labels all work unchanged on podman, confirmed by running them. That is why
this is a two-entry lookup rather than an interface: only the probe differs.
Detected rather than declared, because adoption keeps what the machine already
has (research 012), which hardcoding one runtime contradicts.
A machine with neither now says so, naming both: "docker: command not found" on
a machine deliberately running podman sends the reader after the wrong thing.
Verified end to end against real docker (container created, running, labelled)
and against an empty PATH (refused, naming both runtimes).
Two injections per behaviour, all confirmed to bite. One injection produced a
build failure that my check read as "no bite" for the third time, so the check
now distinguishes them.
|
||
|
|
9a9937b7e6 |
A struct per resource kind, instead of one struct with every field
Jochen asked why we don't simply have dedicated structs. We should, and the
flat struct was me extending an existing pattern rather than questioning it.
Before: one Resource struct carrying path, content, mode, unit, state, package,
image, name, env, ports, volumes, args, command, verify and in. Because a file
and a container shared it, nothing stopped {"type":"file","image":"postgres"},
so a `uses` map listed which fields each kind was allowed to carry -- a second
place to keep current, and the kind nobody updates is the one that silently
accepts a field the host will never read.
Now: Directory, File, Service, Package, Container and Action are separate
structs behind a Resource interface. File has no Image field, so the mistake is
not detected -- it is unrepresentable. Adding a field to a kind is the whole of
adding it; there is nowhere else that has to agree.
Parsing is two passes: read the envelope and each resource's raw bytes, peek at
"type" to choose the struct, then decode into it. Peeking is lenient on purpose
-- reading strictly there would report an unknown field before knowing which
fields are known.
Unknown fields are found by comparing the JSON keys against the struct's own
json tags rather than by catching the decoder's error. The decoder stops at the
first unknown field, and RefusalError promises every problem at once: a caller
fixing one field at a time learns the next only by running again. Caught by
testing the refactor against a real declaration -- a container carrying both
`unit` and `mode` reported only one of them.
apply.go switches on the concrete type instead of a string, so a new kind that
has no applier is a compile error rather than a runtime default branch.
No behaviour change otherwise. All existing tests pass unmodified except two
that reached for fields the interface no longer exposes.
|
||
|
|
337126603e |
Complete the host's vocabulary: package, container, action
The three shapes the substrate bootstrap needs and the host did not have. Until now tier 1 could not be raised at all -- step 0 is a package, step 1 a container, steps 2 and 3 actions -- so every line of the tier 1 and 2 designs was unbuildable. package -- present, never upgraded, never uninstalled. Removal is "forgotten", not "removed": the host cannot know what else needs the package, uninstalling a container runtime because a declaration changed would stop every container on the node, and the machine may have had it before the mesh saw it. Reporting it removed would claim an effect the host declined to have. container -- identified by a label carrying a digest of the declaration that made it. Comparing every field the runtime reports cannot be done reliably: a runtime normalises, defaults and reorders what it is given, and that is indistinguishable from real drift. There is no in-place update; a container's configuration is fixed at creation, so any change is a replacement, and saying so beats a partial update that leaves the running thing half-declared. This is the one shape the host removes, because it is the one the host created. action -- bundle-only, per ADR 0047. Verify is mandatory and does double duty: it is the idempotency check as well as the read-back. The host does not know what a database is, so "is it already there" is a question only the declaration can ask. `in` runs the action inside a named container, which steps 2 and 3 need. Parse now refuses actions; ParseTrusted permits them. The safe path is the default and the permissive one has to be named. The bundle and a local file handed to a root process use ParseTrusted; the link will use Parse. Also replaced the per-type "fields this type ignores" check with a field-set diff stated as what each type USES. The negative form needs every type revisited whenever a field is added, and the one nobody revisits silently accepts a field it will never read. Images must be pinned by digest (ADR 0046). A bundle naming a tag pins nothing. Verified against a real machine, not only fakes: an action ran and was idempotent on the second apply; an action that exits zero and satisfies nothing fails the apply; a real container was created, labelled, replaced when its declaration changed, exec'd into, and removed; a real package query round- tripped. Each new test was also confirmed to fail on an injected fault -- five injections, each breaking exactly its own test. One existing test changed: a vanished unit is now reported "forgotten" rather than "removed", which is what actually happened. |
||
|
|
9d8239afe8 |
Stage 2 — the host applies a declaration
A declaration is JSON, versioned, and an ordered list of resources with stable identities (novox/hq ADR 0043). The vocabulary is directory, file and service, and anything outside it — an unknown version, type or field — refuses the WHOLE declaration. A host that skipped what it did not understand would apply most of what it was sent and report success. It converges rather than executes: applying twice changes nothing the second time, and applying to a drifted machine returns it. A mode is maintained rather than set, because a permission applied at creation is not a permission held — this repository has paid for that once already. It owns a footprint and only that. What it applied and is no longer declared is removed; what it did not create is never touched. Removal runs FIRST, because a resource leaving a declaration while another arrives at the same path is an ordinary rename, and removing afterwards would delete the file just written. The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing can be removed without knowing what was applied. It is written atomically, refuses to start empty when it exists and cannot be read — believing it owns nothing would leave everything behind forever — and is saved even when an apply fails, because what was applied before the failure is on the machine either way. Three faults found by running inside a raised machine rather than by reasoning: A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`, exactly as a stopped one does. So declaring a unit stopped reported success for a unit the host cannot manage at all — absence read as satisfaction, which is 04-ISSUES/007 wearing a different hat. LoadState separates them. Removing an orphaned service whose unit has since been uninstalled failed the whole apply, and a host holding such a record could then apply NOTHING, ever, with no way out but editing its state by hand. Removal is now idempotent for the same reason os.RemoveAll is. And the flag parser was wrong in the same way twice: fixing `mesh-host inventory --json` by taking the subcommand off the front left `mesh-host apply decl.json --dry-run` broken identically, because the standard library stops at the first non-flag argument wherever that argument is. Parsed in a loop now. 30 new tests, 55 in total. |