With a gate on the first machine and a rollback after it, a module's build rolls out by
default. The ones kept back say why: the network path a rollback could not cross, the
providers every consumer on a machine drops with, and the stores holding the photos.
A merge's deleted files are announced, so a module whose manifest went is forgotten
rather than asked to build (the public-acme plan failure).
Both caught a failed write and took the event, losing it silently; the SDK's rule is to throw when
the work was not done. A failed write now throws so the bus offers the event again, and is spooled
on disk at once; on its last delivery the spooled event is taken, and a background pass replays the
spool once writing works. The runtime does not pass the delivery count, so the spool counts failed
deliveries itself, across restarts. Over its bound (1000 events or 30 minutes) the last delivery is
no longer taken, so the bus gives it up and the controller raises max-deliveries - the one existing
condition that names a consumer which cannot keep up - while the spool still holds it.
Writes are idempotent by event: the trail skips an id it already wrote; the usage upsert keeps the
reading observed latest (migration 2), so a late replay never overwrites a newer one. Each module has
a status tool for the spool, declared as valuable data (ADR 0233). model-usage moves to the bundle
shape (ADR 0198) with its schema in a prepare step and numbered migrations; its old container shape
had no image. Both on mesh-sdk 0.1.13.
The log-only handlers of redis, mssql, mosquitto, mongodb, mesh-vault, showcase and the catalogue no
longer throw a TypeError on an event without a body.
Walks run at most daily and stop after ten minutes or two million files; datasets are read from
their counters and large items from their top level only. The database platform's tables are
dumped with pg_dumpall rather than copied as live files.
Backup lines are derived from each module's data section instead of written by hand; the holder
measures declared items, reads the array under them, and deletes a retired item only after a last
restore point; the Go providers say each held consumer's size so an empty replacement is seen.
public-acme ran nothing and had one consumer. The proxy now states the issuer itself, byte for byte
what the binding rendered, so its account directory and every certificate stay put. dhcpcd and
cloudflare-dns are assigned nowhere and nothing requires what they provide.
Five passes are twenty-five seconds, shorter than a controller restart, a store
reconnecting or a file half written; the operator asked for both (hq ADR 0230).
The hourly release of ADR 0229's brake still ended in the mesh acting alone on
a mistake. A consumer now stays active until the same unasked set holds for
five passes, waits for a person past three or half of those held, is disabled
and marked rather than withdrawn, comes back as it was when asked again, and is
deleted only through the provider's delete tool. The backend keeps the mark, so
a restart forgets nothing and finds what was withdrawn before.
Issue 241 withdrew seven consumers in one pass on one misread file; its fix
refuses a file it cannot read, and a file read whole that names nobody
still withdraws everybody. A pass that would withdraw more than one
consumer at once, or more than half of those it holds, now withdraws
nothing: each consumer it kept is announced provisioner.failing with the
class withdrawal-braked, so the controller raises it as a condition and
the operator is told, and one is let go each hour while the mesh goes on
not asking for them. A consumer asked for again is kept and said recovered.
Postgres and keycloak carry the harness identically.
The controller now replaces a given at-start secret after the module's
first good start and on rotate (hq ADR 0228); a key only OpenAI can
issue must say so, or a fresh random value would take its place.
A consumer made on 2026-10-06 was handed three hours of raised-and-cleared
conditions at once and the holder said each as new: 20 desktop notifications
in a second. What is said is now decided by the controller's open set and by
an event's own time, never by its arrival; bursts are one message, the
desktop gets warnings at most every 15 min, the cap is said once, and the
first minute after start says only the urgent conditions still open.
Replays of that morning's 96 events are tests. Also drops two committed
binaries. (novox/hq issue 271)
logflare 1.4.0 copies LOGFLARE_API_KEY into its default user and
POSTGRES_BACKEND_URL into every source's backend once, when it creates
them, and never reads either again. On ace every source still points at
db:5432, which resolves nowhere, so analytics stored no logs; and the key
that leaked into the log before #85 could not be replaced, because the
mesh had to mark it "applied".
The start script now runs logflare's migrations and then reconcile.exs
through `logflare eval`, before logflare starts: it uses logflare's own
Users and Backends contexts to set the default user's key to
LOGFLARE_API_KEY (clearing old_api_key) and to point each postgres
source backend at POSTGRES_BACKEND_URL. It changes nothing that already
matches, says what it changed without printing the key or a password,
and stops the start when it cannot finish.
With the key taken at every start, the secret is "at-start" and `secret
rotate` works. analytics restarts on its env file and studio on its env
file too, so a rotated key reaches every reader.
The mesh noticed 48 core failures in six days and told nobody (ADR 0227).
messenger holds the operator-channel seat: it consumes the controller's
condition events and sends them to Telegram and the desktop notifier,
deduplicated by key, reminded once, edited on clear, capped at 20 an hour
with the rest folded, and refusing anything carrying an address, a path or
a secret. mesh-watcher, on a machine other than the control node, sends to
Telegram directly when the self-check heartbeat or the bus goes silent.
vector handed logflare its API key as ?api_key= in every sink URL, and
logflare 1.4.0 prints a failed request's whole URL in its Plug.Cowboy
error report. Its ingest fails on every request here, so the key was in
the analytics log about every ten seconds. The report is an error, so no
log level hides it.
Every sink now sends the key in the x-api-key header, which logflare
reads first. A start script refuses to start logflare while the vector
config it is given still puts the key in a URL.
The key is marked "applied", not "at-start": logflare writes it into its
default user once and never updates it, so the mesh must not rotate it
by restarting.
Consumers were held to an S3 access key's 20 characters whatever they
required. Each provider now says what its backend keeps: minio 20,
PostgreSQL and MongoDB and DNS 63, Gitea 40, a mailbox 64, SQL Server 128,
Keycloak 255, unbounded where the store has no limit, and none for the
resolver and route provisions, which keep no name of their consumers.
Needs the controller that reads the field (mesh-controller, ADR 0225).
letta printed two passwords into its log for weeks and nothing noticed,
and docker_logs handed them to whoever asked. docker_secrets_in_logs
compares each container's recent lines with the secret-named values of
its environment, the passwords in its URIs, and any URI carrying a
password, and names what it found by container, module and variable -
never the value. docker_logs redacts the same values before answering.
letta 0.6.8 prints LETTA_PG_URI whole (startup.sh, alembic, server.py)
and its server password when it starts in secure mode, so both were in
the container's log on every one of its restarts. Newer letta still
prints both, and neither is a log level.
The URI now names no password: libpq reads it from a mounted pgpass
file (PGPASSFILE). The one print of the server password is rewritten by
a start script before the server starts, and the script refuses to start
letta if that print, or a password in the URI, is still there - a letta
that does not start says why; one that leaks says nothing.
Both own secrets say they are read at start, so `rotate` can replace the
server password the mesh made.
On 2.10.29 such a consumer was moved past a message now and then without
handing it over; the controller's events consumer has seven filters, and a
merge on the stream never reached it. The test reproduces the skip on 2.10.29
and keeps the image's release equal to the server it tests.
The bootstrap once created the temporary admin and then failed on a held port; marked only after
it succeeded, the cleanup did not know the admin existed and left it.
Twice the identity provider's admin kept an older password than the one the
mesh minted (an adopted, then a moved database), and the provisioner failed
every consumer until it was repaired by hand (hq issue 179). The module now
checks the admin's login and repairs a refusal itself through the server's
bootstrap command, verifies, brakes a failed repair and announces it, and
stops asking the server while refused. Ported to Go to change it.
A provider failed every consumer for a day and said so only in its journal
(hq issue 179). The provisioner loop now emits provisioner.failing after five
minutes without a success — create, check or secret — and repeats it every
fifteen; provisioner.recovered on the next success, on withdrawal, and on the
first success after a restart, so the controller can name it in status
(hq ADR 0224).
Requiring wildcard-resolution gave each a consumer identity; mesh_<machine>_networkmanager is
over the 20 characters a backend keeps, and the anchor's declaration, which carries every
consumer's grant, could not be composed.
Two files say one fact, the machine's name, and nothing owned /etc/hostname.
The name written is the operator's hostname setting, with no default: three
of four machines call themselves something other than their mesh name, and
renaming one is the operator's call. It takes effect at the next boot.
The program that manages a machine's network is the one that would rewrite
the resolver file, so its module now writes it: networkmanager,
systemd-networkd and dhcpcd render the same template from the resolver's
holders. resolv-conf declares nothing for one release, so every machine
hands the file over in one apply; it is removed once unassigned everywhere.
letta crash-loops on 'type "vector" does not exist': pgvector is not a
trusted extension, so only the provider's superuser can create it, and
the provisioner never did. A contribution may now name extensions; the
provider creates each (IF NOT EXISTS, available ones only) in the
consumer's database on every pass. Go per the standing rule for a
TypeScript module that changes. letta asks for vector.
Next.js binds the address HOSTNAME names; docker sets HOSTNAME to the container's id, so studio
answered only on its network address while its image's health check asks localhost, and it read
unhealthy while working. HOSTNAME=:: as the upstream compose file sets it.
musl asks every nameserver at once and takes the first reply, so a public
resolver's NXDOMAIN for a mesh name beat the mesh's answer in every Alpine
container (hq ADR 0223). resolv-conf now renders /etc/resolv.conf from the
holders of mesh-dns-resolver, this machine first when it holds one; dnsmasq's
comments say the seat may have several holders.