On 2026-10-07 adopting its images' health checks recreated nine of mail's
containers in one send and the operator's phone could not reach the mail.
A module people use directly is moved at a moment a person chooses.
Seven modules' images ship a check the mesh never read. Adopted by name where
it says healthy on the live mesh today: nine of mail's containers (not its
antivirus, whose six-minute start is past the five-minute bound, nor its cache,
whose image ships none), the certificate authority, the spreadsheet app, four
of the database suite's (the studio among them, with the address it binds
fixed), the flow editor and the chat client. And the endpoints four services
already declare, looked at from the machine: tcp on the database, the cache,
the document store and the broker; http on the website and the dashboards.
The count of undeclared falls from 93 to 70.
A field that is optional for ever is one half the catalogue never gets. The
catalogue's merge check now holds the controller's count of long-running
resources that do not say how they are ready to the number kept in
health-undeclared: a change that raises it fails, one that lowers it must
write the new number. Until the controller the mesh runs counts (Phase B), the
check says it did not count.
Branch protection requires mesh/merge-gate (and mesh/repo-check on the core
repositories) with no admin override, and the forge combines warning as a
failure. A note is now a success that says it; a repository without a
merge-check.sh is a success where repo-check is not required and a failure
for a person where it is; the status tool refuses the merge check's contexts.
A pull request's head announced again from the bus's history at
mesh-delivery's first start was proposed with no verdict and nothing
ever asked its check: the controller had taken that announcement long
before, and stalled raised it after an hour for the operator.
A proposed delivery with no verdict and no check asked is now asked
through delivery-check once it has waited past a grace longer than a
check takes; an announcement carrying its head's decided gate status
takes it. The proposed bound runs from the ask, and H2's close may
re-ask once (a table row) before the delivery is the operator's.
mesh/merge-gate pass: builds docker, keycloak, minio, mosquitto → ace, g14, novox, shanks; no bus step; 2 wait(s) for a person; every machine composes with…
mosquitto passed the broker's admin password to mosquitto_ctrl as -P on
every docker exec, and the container runtime keeps every exec's command
line in its event stream, where docker_events returned it. The admin
credentials now reach mosquitto_ctrl as a 0600 options file fed on
stdin, client passwords at its own prompt, and an argv carrying a secret
is refused before it runs. The admin secret says it is taken at start:
the bootstrap re-runs when the mesh replaces it and re-keys the broker
online from the value it last applied, so it can be rotated.
docker_events redacts what an exec's command line carried, and
docker_secrets_in_events names such secrets by name. keycloak's repair
hands kcadm its passwords through KC_CLI_PASSWORD; minio gives mc its
root alias through MC_HOST_mesh.
One module answers 'did my change go out' for a commit and orders a
cross-repository change: one compiled state table, its state on the bus,
every transition said, noted on the commit and shown on the pull request.
The forge's holder gains the note, view and status tools it asks with, and
says closed pull requests and a merge's head and statuses.
The gate's status says what the change does and how it was judged; a change that builds
something gets its plan as a comment — the tiers, what each machine receives, and what
is not an ordinary send.
The controller asks its planner what a pull request reaches, as if merged: a module whose
manifest the change deletes is one the merge removes, and the plan's dependents are said
on the gate's status beside the modules it moves.
The controller now decides what a pull request's check runs from the mesh's module
graph, so the announcement carries the changed directories that hold a module at the
head and whether the head has a merge-check.sh. A verdict sets mesh/merge-gate (the
gate, with the modules it judged) and mesh/repo-check (the repository's own tests).
gitea_branch_protection_get/set let the operator's agent make those statuses required.
The catalogue's merge-check.sh leaves the gate to the build seat and keeps its own layer:
every manifest through module check, and the touched Go modules' tests.
The forge's announcer announces each new head of an open pull request as pull.updated, once,
and marks the head pending; the controller asks the build seat to check it against every
machine of the mesh's facts, and says the verdict as checked, which the forge's holder sets as
the head commit's status mesh/merge-gate - an error never as a success - with the check's own
account as a comment when it is not a pass. merge-check.sh is the catalogue's check: every
manifest through the running controller's module check and merge gate, and the Go tests of each
module the change touches, a module whose dependencies cannot be fetched said as not tested.
The controller read a changed file as shared code unless its directory was a module it holds or
the merge also changed that directory's manifest. A merge touching modules/showcase/index.ts -
the catalogue's reference module, held by no machine - therefore rebuilt all 103 modules built
from this repository on 2026-10-06, 88 of them byte-identical, with the build agent first only
because everything else is built by it.
Whether a directory is a module is a fact of the repository at the commit, so the announcer now
looks it up: every directory above a changed file (never the root) is asked for its module.json
at the merge commit, and the ones that have one go out as module_dirs, with module_dirs_said.
Not said when the file list is cut, past 300 directories, or when the forge cannot be asked;
the controller then keeps its old rule, which rebuilds too much rather than too little.
modules/showcase stays: TestTheShowcaseModuleIsAValidManifest in mesh-controller parses it and
hq to-be 18 and 20 name it as the reference module.
With a gate on the first machine and a rollback after it, a module's build rolls out by
default. The ones kept back say why: the network path a rollback could not cross, the
providers every consumer on a machine drops with, and the stores holding the photos.
A merge's deleted files are announced, so a module whose manifest went is forgotten
rather than asked to build (the public-acme plan failure).
Both caught a failed write and took the event, losing it silently; the SDK's rule is to throw when
the work was not done. A failed write now throws so the bus offers the event again, and is spooled
on disk at once; on its last delivery the spooled event is taken, and a background pass replays the
spool once writing works. The runtime does not pass the delivery count, so the spool counts failed
deliveries itself, across restarts. Over its bound (1000 events or 30 minutes) the last delivery is
no longer taken, so the bus gives it up and the controller raises max-deliveries - the one existing
condition that names a consumer which cannot keep up - while the spool still holds it.
Writes are idempotent by event: the trail skips an id it already wrote; the usage upsert keeps the
reading observed latest (migration 2), so a late replay never overwrites a newer one. Each module has
a status tool for the spool, declared as valuable data (ADR 0233). model-usage moves to the bundle
shape (ADR 0198) with its schema in a prepare step and numbered migrations; its old container shape
had no image. Both on mesh-sdk 0.1.13.
The log-only handlers of redis, mssql, mosquitto, mongodb, mesh-vault, showcase and the catalogue no
longer throw a TypeError on an event without a body.
The restic holder copied JetStream's store while the server wrote it; such a
copy may not restore. The nats image now carries mesh-nats-snapshot, run by
the declared dump under the module's own bus account (snapshot API only):
every stream one at a time, flow-controlled, into one tar with a manifest of
counts, sequences and checksums. Restore builds a new store beside the live
one with the bus's own server; a person swaps it in. Proven against
throwaway nats 2.11 servers being written to during the snapshot.
Walks run at most daily and stop after ten minutes or two million files; datasets are read from
their counters and large items from their top level only. The database platform's tables are
dumped with pg_dumpall rather than copied as live files.
Backup lines are derived from each module's data section instead of written by hand; the holder
measures declared items, reads the array under them, and deletes a retired item only after a last
restore point; the Go providers say each held consumer's size so an empty replacement is seen.
public-acme ran nothing and had one consumer. The proxy now states the issuer itself, byte for byte
what the binding rendered, so its account directory and every certificate stay put. dhcpcd and
cloudflare-dns are assigned nowhere and nothing requires what they provide.
Five passes are twenty-five seconds, shorter than a controller restart, a store
reconnecting or a file half written; the operator asked for both (hq ADR 0230).
The hourly release of ADR 0229's brake still ended in the mesh acting alone on
a mistake. A consumer now stays active until the same unasked set holds for
five passes, waits for a person past three or half of those held, is disabled
and marked rather than withdrawn, comes back as it was when asked again, and is
deleted only through the provider's delete tool. The backend keeps the mark, so
a restart forgets nothing and finds what was withdrawn before.
Issue 241 withdrew seven consumers in one pass on one misread file; its fix
refuses a file it cannot read, and a file read whole that names nobody
still withdraws everybody. A pass that would withdraw more than one
consumer at once, or more than half of those it holds, now withdraws
nothing: each consumer it kept is announced provisioner.failing with the
class withdrawal-braked, so the controller raises it as a condition and
the operator is told, and one is let go each hour while the mesh goes on
not asking for them. A consumer asked for again is kept and said recovered.
Postgres and keycloak carry the harness identically.
The controller now replaces a given at-start secret after the module's
first good start and on rotate (hq ADR 0228); a key only OpenAI can
issue must say so, or a fresh random value would take its place.
A consumer made on 2026-10-06 was handed three hours of raised-and-cleared
conditions at once and the holder said each as new: 20 desktop notifications
in a second. What is said is now decided by the controller's open set and by
an event's own time, never by its arrival; bursts are one message, the
desktop gets warnings at most every 15 min, the cap is said once, and the
first minute after start says only the urgent conditions still open.
Replays of that morning's 96 events are tests. Also drops two committed
binaries. (novox/hq issue 271)
logflare 1.4.0 copies LOGFLARE_API_KEY into its default user and
POSTGRES_BACKEND_URL into every source's backend once, when it creates
them, and never reads either again. On ace every source still points at
db:5432, which resolves nowhere, so analytics stored no logs; and the key
that leaked into the log before #85 could not be replaced, because the
mesh had to mark it "applied".
The start script now runs logflare's migrations and then reconcile.exs
through `logflare eval`, before logflare starts: it uses logflare's own
Users and Backends contexts to set the default user's key to
LOGFLARE_API_KEY (clearing old_api_key) and to point each postgres
source backend at POSTGRES_BACKEND_URL. It changes nothing that already
matches, says what it changed without printing the key or a password,
and stops the start when it cannot finish.
With the key taken at every start, the secret is "at-start" and `secret
rotate` works. analytics restarts on its env file and studio on its env
file too, so a rotated key reaches every reader.
The mesh noticed 48 core failures in six days and told nobody (ADR 0227).
messenger holds the operator-channel seat: it consumes the controller's
condition events and sends them to Telegram and the desktop notifier,
deduplicated by key, reminded once, edited on clear, capped at 20 an hour
with the rest folded, and refusing anything carrying an address, a path or
a secret. mesh-watcher, on a machine other than the control node, sends to
Telegram directly when the self-check heartbeat or the bus goes silent.
vector handed logflare its API key as ?api_key= in every sink URL, and
logflare 1.4.0 prints a failed request's whole URL in its Plug.Cowboy
error report. Its ingest fails on every request here, so the key was in
the analytics log about every ten seconds. The report is an error, so no
log level hides it.
Every sink now sends the key in the x-api-key header, which logflare
reads first. A start script refuses to start logflare while the vector
config it is given still puts the key in a URL.
The key is marked "applied", not "at-start": logflare writes it into its
default user once and never updates it, so the mesh must not rotate it
by restarting.