The empty-resources guard refused every empty body as a likely mistake,
with no way to say emptiness was meant — so the control plane could
never tell a node to drop its last resource. The envelope gains
owns_nothing: with it, an empty declaration is applied (the node drops
what the mesh owned); without it, empty is still refused, so a
truncated or mis-composed body cannot silently strip a machine. One
test, both directions.
The host parses MTU from the found [Interface] and reports it, so the
mesh's interface can come up with the same MTU when it takes the tunnel
over. A path tuned to 1380 regresses to the 1420 default otherwise —
invisible to ping, fatal to TLS handshakes and transfers over that path
(novox/hq: the mesh had no MTU concept). Zero when the config named
none, and the mesh writes no MTU line then.
dns and ip were declared, validated, handed to the runtime — and part of
no comparison, so their first deployment compared every container equal
and changed nothing, silently. The same shape as 04-ISSUES/045: a field
that is not in the spec is a field that can never reach a container that
already runs.
Mailu's 2024.06 admin refuses to serve behind a resolver that does not
validate DNSSEC, and the runtime's own forwarder (127.0.0.11) validates
nothing — so a module shipping its own validating resolver had a
resolver nothing could be pointed at. Found live, blocking a cutover:
the admin sat unhealthy, submission answered 454, and the declaration
language had no words for the fix.
Two fields on a container, both handed to the runtime verbatim: dns —
the resolvers it asks — and ip, its static address on its user-defined
network, which exists for exactly one shape: a container others must
reach before name resolution works, the resolver itself being the case
that forced it. Both take only addresses and are refused on arrival
otherwise — a name here would reach the runtime verbatim and be refused
at create, after the old container was already gone.
docker inspect <name> resolves across every object kind, not just
containers. A module regularly names a network the same as the
container that joins it (keycloak does this today, ordinarily) — so
when the container does not exist yet but the same-named network
already does, the bare form answers with the network's JSON instead
of reporting the container absent, and the template these callers use
(.State.Running) fails to execute against it entirely.
Live on novox tonight: minio's LB container, named the same as its
network ("minio"), could never be created — every apply crashed on
"the container runtime could not say whether minio is here", stuck
since first push, because the check itself never got a clean answer.
Fixed at every call site asking a container's state by name
(containerState, inspectFound, NamesFree, raiseGiteaServer,
containerRunning) by scoping to `docker container inspect`, matching
the type-scoped form this codebase already uses correctly for
networks, volumes and images elsewhere. Also scoped the one image
inspect that was still bare (publish.go), for the same reason.
mesh-host runs as a host-level service (nox-mesh-host.service), not a
Docker module — merging this does not redeploy it. The live novox
failure persists until the service itself is rebuilt and updated.
Review of the ADR 0105 build (hq ADR 0105). The takeover stopped the found
unit and then found out whether the mesh's interface would do; a start that
failed left the machine with no tunnel at all.
Now nothing is stopped until the declared interface listens on the found port
at the found address and the key file it names holds the found key — the
refusal names the remedy — and a mesh interface that fails to start after the
takeover has the found unit started again, with the account saying so. The
account has three states (not taken, taken, down) and is given on every
takeover, failure included. An interface raised by hand is looked at again
for a moment and then refused naming `wg-quick down`. A found unit started
again by hand beside the mesh's is said, not stopped: on the hub it cannot
hold the port, and on a spoke two interfaces with one key would fight.
`mesh-host overlay take --tunnel <iface>` is the path for a node that
enrolled before the mesh knew to take a tunnel over: the found key becomes its
overlay key — identity, sealing and serving keys untouched, so nothing sealed
to the node is remade — and the mesh is told with a rekey signed by the
identity key, over the key left, the key taken and the tunnel. Told first,
written second, so a run again puts right whichever half did not happen.
The plan says what an apply would change from the declaration and the
record, before the machine is touched. The apply now recreates a container
when the content of a file it reads at creation changed, and the plan said
"check" for every recorded container — true, but a preview that hides the
one step somebody asked about.
So a recorded container whose record of what it read differs from what this
apply will hand it — a plain file declared here, by its declared content;
otherwise what this host last wrote at that path — is planned as an update
naming the file, the same comparison applyContainer makes. What the record
cannot settle stays a check: a file neither declared nor recorded is read
from the machine by the apply, not by the plan; and a container with no
record of what it read was labelled before the host kept that record and is
accepted as it is.
novox/hq 04-ISSUES/103, 104
Review of the first cut found four things.
A directory mounted into a container is no longer looked inside, not even
for the files this host wrote there. The controller records every
provider's received and contributions file as a plain file under a mounted
directory, so folding those in would have recreated the route proxy — which
re-reads its routes live, by design — on every route change, and killed
every provisioner sidecar, which polls what it receives, mid-reconcile on
every grant. Whether a service reads a file under its directory once or
watches it is the service's; restart-on is how a module says "once", and it
stays the opt-in. Env-files and files mounted directly remain by content.
Genesis wrote the superuser secret as `value\n`; `secret accept` strips the
line ending by design, so the postgres module declared `value` — and with
a mounted file's content in the spec, phase three would have recreated the
store it meant to adopt in place, with the temporary control plane
connected to it. Genesis now writes the value alone. readCredentialFile
tolerated both endings already. Pinned with the bytes the genesis code
path writes, then the module's declaration of the same container: it must
reconcile.
A container carrying a label from before the host folded in what it reads
is accepted rather than recreated, when that label matches the spec as it
used to be computed: what it reads is recorded then, a change is caught
from that record from the next apply on, and the label is renewed at the
next genuine recreate. Recreating them all would have been a restart storm
across the mesh in declaration order, the store first. The trade-off is
stated in the code: a container already stale at upgrade time is not
caught, and could not have been either way.
The record of what a container read is looked up by its name when its
declared id has none — the bundle's `store` becomes `postgres.server` for
the same container — so a change on the day it is adopted still names the
file. The by-target lookup takes the most recently applied record, since
the bundle's record for the same target is never removed by the mesh's.
novox/hq 04-ISSUES/103
The host decided whether a container was still the one declared by a digest
of its declaration, and the declaration names an env-file's path and a
mount's path — never what is in them. So when the store was given a new
port, the host rewrote the forge's and the analytics service's environment
files, correctly, and left both containers running with the old port in
their environment: a container reads its env-file when it is CREATED, and
`docker restart` hands it the same environment again. Both looked healthy
until they answered 502.
What a running container takes in at creation is now part of its spec, by
content: every env-file, a file bind-mounted into it, and every file this
host wrote at or under a directory bind-mounted into it — the secrets,
bindings and configs under a module's state directories. The digest is the
one the store already records for a file the host wrote (`wrote`), read
from the state as it stands when the container is reached, so a file
rewritten earlier in the same apply is already the new one; a file the host
has no record of — an env-file a predecessor left, the superuser secret
genesis writes before any declaration names it — is read from disk, which
is what keeps adopting a running store in place a reconcile and not a
recreate.
Deliberately not part of it: what else is in a bind-mounted directory,
which is the service's own data and changes while it runs; a named volume;
a seed created once, which digests as the seed the host wrote and not as
what has grown in it; and a step — a run-once or scheduled container reads
its files when it runs and runs fresh each time. On an adopted node a held
container is held before any of this is looked at.
The host records what each container was created reading, per file, so
the recreate can say which file changed — "recreated: <file> changed" in
the report and, now with its detail, in the log. A container made before
this record existed is recreated once and says so.
novox/hq 04-ISSUES/103
Review of the fix for hq issue 104 found three faults in it. A file applied
on an enrolled node — the mesh's own last declaration included — is applied
as the bundle is, so its resources are recorded as the machine's own and
what the mesh declared reads as undeclared: the plan removed the foundation.
`apply FILE` is for a machine the mesh has not spoken to, and is now refused
saying so whenever declared.json exists. The plan looked at what is held
before what the declaration says is taken, so the one cutover ADR 0100 says
must be previewed read as a hold; it now decides in holdOnAdopted's order,
models a step run inside a held container, and a test holds the plan's
sequence to the apply's outcomes. Genesis wrote the mode on every run, so a
re-run after `converge` left the state saying adopted while the kept,
signed declaration said converged, and the reconcile loop refused every five
minutes with no delivery coming to end it: genesis now writes the mode only
when none is recorded, and where the state and the verified kept declaration
disagree, the kept declaration wins and the repair is said.
Also: a file lock beside the state, taken by the link service, the host's
own commands and the installer alike, so a `reconcile` run by hand no
longer races the loop's save — chosen over refusing while a named service is
active, which would miss a `mesh-host run` started by hand; `--json
--dry-run` emits {plan} like an apply emits {plan, report}; the README's
duplicate flag line; and the bundle refusal is about the digest, not a claim
the carried bytes can never match what genesis applied.
On an adopted machine the private network takes the predecessor's tunnel
over in place (hq ADR 0105). Genesis finds the one interface up besides the
mesh's own, settles the hub's port and the mesh's range on it, and skips
ADR 0100's non-overlap check for a range that is now the tunnel's; a
--hub-port or --overlay-range that disagrees is refused naming the tunnel's.
At enrolment the found interface's private key becomes this node's overlay
key — the one credential the mesh takes rather than mints — stored where a
generated one is stored, never printed and never sent; the tunnel (port,
address, range, peers) travels with the keys so the mesh composes from it
before the first declaration.
The interface's service may say what it takes over. Before the mesh's unit
starts, the found configuration is kept like any held file and the found
unit is stopped and disabled; nothing is flushed, and an interface still up
after its unit stopped refuses the takeover rather than half-working. The
report says what was carried: interface, port, range, peer count, taken or
not, and where the original was kept.
An operator ran `mesh-host reconcile` on an adopted control-node with twelve
modules assigned. It applied the bundle the host carries — the genesis
declaration, foundation only, converged: recreated the store, failed on the
broker's held port, wrote the converged base filter and started its service,
and stopped at the first failing action. The filter closed the machine for
forty-five minutes. The host reported the node adopted in every report, the
declaration said converged, and nothing compared the two; nothing was printed
before acting (hq issue 104).
The host now records the node's mode — from every declaration the mesh sends,
and at genesis from what the operator said — and refuses, at the point of
application, a declaration that says the other mode, naming both and the act
that changes it. Only a declaration the link delivers, signed, changes the
mode: that is how `converge` and `adopt` arrive, so the flip still works and
nothing else can do it. Genesis marks the bundle consumed, with the digest of
what it applied, so `reconcile` holds a node the mesh has spoken to against
what the mesh last said and never the bundle, and refuses the carried bytes
when they are not what genesis applied. A file is refused when it is not what
the mesh last said: a declaration carries no sequence and no issued-at, so the
host cannot tell older from newer, and says so. Both commands print what they
would change — a hold, a removal, an action named as one — before touching
anything, and --dry-run is that list and nothing more.
Registering a module is an overwrite. Recording the forge's port ran `module add`
every time, so a genesis re-run pointed at an older catalogue would replace the
manifest of a forge that is built and assigned — with a push a few lines later.
Registering is only here so a settings row has a module row to hang on, and that
row is already there on a mesh that knows the forge. So: ask first, and skip.
Two comments narrowed to what is true. What follows the node's setting is what
the mesh derives from a module's ports — its container mapping, its filter rule,
its opening and what it serves. The forge's own address in its runtime's
environment (hq 088) and its route contribution's port do not, and are already
wrong for any port the mesh assigned. And a settings layer is the module's, not
one resource's: a second mergeable file on the builder would be given `serves`
too.
novox/hq 04-ISSUES/085
Every foundation port given at genesis became a per-node setting of the module
that binds it, except the package registry's: that one was fixed by rewriting
the builder's manifest when the installer registered it. Registering the builder
again from the catalogue undid it, and the forge's own module, when it took the
bootstrap forge over, came up on the catalogue's port — which on a machine where
a predecessor holds 3000 points the builder at the predecessor's forge.
So the rewrite is gone, and the port is recorded twice as a setting, both from
the one input:
- the forge's module is registered at genesis — not assigned, nothing of it runs
— so the controller has something to hold `{"ports": {"3000": <given>}}`
against. Assigning the forge later raises it on the port this machine was
given, and its container, its filter rule, its opening, what it serves and
what consumers are told all read it from there.
- the builder is given `{"serves": {"port": <given>}}`, which merges into the
binding it carries in place of one nothing can resolve yet.
A genesis on the catalogue's port records nothing and registers nothing, so it
does exactly what it did before.
novox/hq 04-ISSUES/085, ADR 0100