Commit Graph
636 Commits
Author SHA1 Message Date
mesh-admin 9d407248b3 Merge pull request 'Back up the bus by the server's own snapshot of each stream, not its live files (hq ADR 0235)' (#93) from feat/bus-snapshot into main 2026-10-06 17:33:13 +00:00
jochen 1fd2914ae1 Say which modules wait for a person's push, and announce the files a merge deleted (hq ADR 0236)
With a gate on the first machine and a rollback after it, a module's build rolls out by
default. The ones kept back say why: the network path a rollback could not cross, the
providers every consumer on a machine drops with, and the stores holding the photos.
A merge's deleted files are announced, so a module whose manifest went is forgotten
rather than asked to build (the public-acme plan failure).
2026-10-06 18:39:13 +02:00
jochen 08265a70ca audit-logger, model-usage: retry a failed write and never lose the event (hq issue 276)
Both caught a failed write and took the event, losing it silently; the SDK's rule is to throw when
the work was not done. A failed write now throws so the bus offers the event again, and is spooled
on disk at once; on its last delivery the spooled event is taken, and a background pass replays the
spool once writing works. The runtime does not pass the delivery count, so the spool counts failed
deliveries itself, across restarts. Over its bound (1000 events or 30 minutes) the last delivery is
no longer taken, so the bus gives it up and the controller raises max-deliveries - the one existing
condition that names a consumer which cannot keep up - while the spool still holds it.

Writes are idempotent by event: the trail skips an id it already wrote; the usage upsert keeps the
reading observed latest (migration 2), so a late replay never overwrites a newer one. Each module has
a status tool for the spool, declared as valuable data (ADR 0233). model-usage moves to the bundle
shape (ADR 0198) with its schema in a prepare step and numbered migrations; its old container shape
had no image. Both on mesh-sdk 0.1.13.

The log-only handlers of redis, mssql, mosquitto, mongodb, mesh-vault, showcase and the catalogue no
longer throw a TypeError on an event without a body.
2026-10-06 18:36:16 +02:00
jochen 419e82cded Back up the bus by the server's own snapshot of each stream, not its live files (hq ADR 0235)
The restic holder copied JetStream's store while the server wrote it; such a
copy may not restore. The nats image now carries mesh-nats-snapshot, run by
the declared dump under the module's own bus account (snapshot API only):
every stream one at a time, flow-controlled, into one tar with a manifest of
counts, sequences and checksums. Restore builds a new store beside the live
one with the bus's own server; a person swaps it in. Proven against
throwaway nats 2.11 servers being written to during the snapshot.
2026-10-06 18:20:51 +02:00
jochen 7524390cad Bound the holder's measuring; dump the database platform (hq ADR 0233)
Walks run at most daily and stop after ten minutes or two million files; datasets are read from
their counters and large items from their top level only. The database platform's tables are
dumped with pg_dumpall rather than copied as live files.
2026-10-06 17:00:23 +02:00
jochen 685cb1cb1b Declare every module's data; the backup holder measures it (hq ADR 0233)
Backup lines are derived from each module's data section instead of written by hand; the holder
measures declared items, reads the array under them, and deletes a retired item only after a last
restore point; the Go providers say each held consumer's size so an empty replacement is seen.
2026-10-06 16:47:49 +02:00
jochen 57d524f4ef Fold public-acme into route-proxy; drop dhcpcd and cloudflare-dns (hq ADR 0226)
public-acme ran nothing and had one consumer. The proxy now states the issuer itself, byte for byte
what the binding rendered, so its account directory and every certificate stay put. dhcpcd and
cloudflare-dns are assigned nowhere and nothing requires what they provide.
2026-10-06 14:59:43 +02:00
jochen 2c15074a0b Retire only once the same answer has held ten minutes as well as five passes
Five passes are twenty-five seconds, shorter than a controller restart, a store
reconnecting or a file half written; the operator asked for both (hq ADR 0230).
2026-10-06 14:28:21 +02:00
jochen e68ef88333 Retire a consumer the mesh stops asking for, and delete only on a person's word (hq ADR 0230)
The hourly release of ADR 0229's brake still ended in the mesh acting alone on
a mistake. A consumer now stays active until the same unasked set holds for
five passes, waits for a person past three or half of those held, is disabled
and marked rather than withdrawn, comes back as it was when asked again, and is
deleted only through the provider's delete tool. The backend keeps the mark, so
a restart forgets nothing and finds what was withdrawn before.
2026-10-06 13:54:10 +02:00
jochen 4238dd8616 Name the SDK harness the Go loop follows: 0.1.11, with the withdrawal brake 2026-10-06 12:29:21 +02:00
jochen 2c151e7c21 Brake a withdrawal larger than its bound (hq to-be 45 Phase 2, ADR 0227 rule 4)
Issue 241 withdrew seven consumers in one pass on one misread file; its fix
refuses a file it cannot read, and a file read whole that names nobody
still withdraws everybody. A pass that would withdraw more than one
consumer at once, or more than half of those it holds, now withdraws
nothing: each consumer it kept is announced provisioner.failing with the
class withdrawal-braked, so the controller raises it as a condition and
the operator is told, and one is let go each hour while the mesh goes on
not asking for them. A consumer asked for again is kept and said recovered.
Postgres and keycloak carry the harness identically.
2026-10-06 12:29:21 +02:00
mesh-admin 12569ed085 Merge pull request 'Mark the OpenAI keys letta and supabase hold as issued outside the mesh' (#89) from feat/a-given-secret-lives-until-the-first-good-start into main 2026-10-06 10:27:21 +00:00
jochen b1bd5c861e Mark the OpenAI keys letta and supabase hold as issued outside the mesh
The controller now replaces a given at-start secret after the module's
first good start and on rotate (hq ADR 0228); a key only OpenAI can
issue must say so, or a fresh random value would take its place.
2026-10-06 12:13:57 +02:00
jochen a4b92c1306 messenger: never say history as news — read the open conditions first, coalesce bursts
A consumer made on 2026-10-06 was handed three hours of raised-and-cleared
conditions at once and the holder said each as new: 20 desktop notifications
in a second. What is said is now decided by the controller's open set and by
an event's own time, never by its arrival; bursts are one message, the
desktop gets warnings at most every 15 min, the cap is said once, and the
first minute after start says only the urgent conditions still open.
Replays of that morning's 96 events are tests. Also drops two committed
binaries. (novox/hq issue 271)
2026-10-06 12:03:19 +02:00
jochen 908d45864a supabase: make logflare's stored key and backends follow its environment
logflare 1.4.0 copies LOGFLARE_API_KEY into its default user and
POSTGRES_BACKEND_URL into every source's backend once, when it creates
them, and never reads either again. On ace every source still points at
db:5432, which resolves nowhere, so analytics stored no logs; and the key
that leaked into the log before #85 could not be replaced, because the
mesh had to mark it "applied".

The start script now runs logflare's migrations and then reconcile.exs
through `logflare eval`, before logflare starts: it uses logflare's own
Users and Backends contexts to set the default user's key to
LOGFLARE_API_KEY (clearing old_api_key) and to point each postgres
source backend at POSTGRES_BACKEND_URL. It changes nothing that already
matches, says what it changed without printing the key or a password,
and stops the start when it cannot finish.

With the key taken at every start, the secret is "at-start" and `secret
rotate` works. analytics restarts on its env file and studio on its env
file too, so a rotated key reaches every reader.
2026-10-06 11:48:57 +02:00
jochen 0ae7933d54 Tell the operator what the mesh finds wrong, and watch the watcher (hq to-be 45 phase 1)
The mesh noticed 48 core failures in six days and told nobody (ADR 0227).
messenger holds the operator-channel seat: it consumes the controller's
condition events and sends them to Telegram and the desktop notifier,
deduplicated by key, reminded once, edited on clear, capped at 20 an hour
with the rest folded, and refusing anything carrying an address, a path or
a secret. mesh-watcher, on a machine other than the control node, sends to
Telegram directly when the self-check heartbeat or the bus goes silent.
2026-10-06 09:35:55 +02:00
mesh-admin 495bb88111 Merge pull request 'supabase: keep logflare's API key out of its log (hq issue 268)' (#85) from fix/supabase-logflare-key-out-of-logs into main 2026-10-06 00:34:13 +00:00
jochen b87dc29706 supabase: keep logflare's API key out of its log (hq issue 268)
vector handed logflare its API key as ?api_key= in every sink URL, and
logflare 1.4.0 prints a failed request's whole URL in its Plug.Cowboy
error report. Its ingest fails on every request here, so the key was in
the analytics log about every ten seconds. The report is an error, so no
log level hides it.

Every sink now sends the key in the x-api-key header, which logflare
reads first. A start script refuses to start logflare while the vector
config it is given still puts the key in a URL.

The key is marked "applied", not "at-start": logflare writes it into its
default user once and never updates it, so the mesh must not rotate it
by restarting.
2026-10-06 02:33:32 +02:00
mesh-admin e5cb7b071c Merge pull request 'State each provision's identity bound (hq issue 263, ADR 0225)' (#83) from fix/263-identity-bounds-per-provision into main 2026-10-06 00:29:56 +00:00
jochen 4769dadf63 State each provision's identity bound (hq issue 263)
Consumers were held to an S3 access key's 20 characters whatever they
required. Each provider now says what its backend keeps: minio 20,
PostgreSQL and MongoDB and DNS 63, Gitea 40, a mailbox 64, SQL Server 128,
Keycloak 255, unbounded where the store has no limit, and none for the
resolver and route provisions, which keep no name of their consumers.
Needs the controller that reads the field (mesh-controller, ADR 0225).
2026-10-06 02:16:27 +02:00
jochen bbb67e41a0 docker: find and hide secrets a container printed into its log (hq issue 268)
letta printed two passwords into its log for weeks and nothing noticed,
and docker_logs handed them to whoever asked. docker_secrets_in_logs
compares each container's recent lines with the secret-named values of
its environment, the passwords in its URIs, and any URI carrying a
password, and names what it found by container, module and variable -
never the value. docker_logs redacts the same values before answering.
2026-10-06 02:13:42 +02:00
jochen f9f27d4878 letta: keep its database and server passwords out of its log (hq issue 268)
letta 0.6.8 prints LETTA_PG_URI whole (startup.sh, alembic, server.py)
and its server password when it starts in secure mode, so both were in
the container's log on every one of its restarts. Newer letta still
prints both, and neither is a log level.

The URI now names no password: libpq reads it from a mounted pgpass
file (PGPASSFILE). The one print of the server password is rewritten by
a start script before the server starts, and the script refuses to start
letta if that print, or a password in the URI, is still there - a letta
that does not start says why; one that leaks says nothing.

Both own secrets say they are read at start, so `rotate` can replace the
server password the mesh made.
2026-10-06 02:13:42 +02:00
jochen 0275c2eeac nats: run 2.11.17, which hands a consumer with several filters every message (hq issue 266)
On 2.10.29 such a consumer was moved past a message now and then without
handing it over; the controller's events consumer has seven filters, and a
merge on the stream never reached it. The test reproduces the skip on 2.10.29
and keeps the image's release equal to the server it tests.
2026-10-06 01:29:20 +02:00
mesh-admin 78328d4ab2 Merge pull request 'keycloak: port to Go and repair a refused admin; providers announce a failing consumer' (#80) from feat/identity-provider-admin-safety-nets into main 2026-10-05 22:39:02 +00:00
jochen 48d4188927 keycloak repair: remove the temporary admin even when the bootstrap failed after making it
The bootstrap once created the temporary admin and then failed on a held port; marked only after
it succeeded, the cleanup did not know the admin existed and left it.
2026-10-06 00:17:35 +02:00
mesh-admin 5c2157b81c Merge pull request 'Rename hosts to hostname, which also writes /etc/hostname (hq ADR 0223 part 3)' (#78) from hostname-module into main 2026-10-05 22:15:29 +00:00
jochen 77fb1ecfb2 keycloak: port to Go and repair an admin that refuses the mesh's secret
Twice the identity provider's admin kept an older password than the one the
mesh minted (an adopted, then a moved database), and the provisioner failed
every consumer until it was repaired by hand (hq issue 179). The module now
checks the admin's login and repairs a refusal itself through the server's
bootstrap command, verifies, brakes a failed repair and announces it, and
stops asking the server while refused. Ported to Go to change it.
2026-10-06 00:13:42 +02:00
jochen 6f1e2f5a0d postgres: announce a consumer failed for minutes, and its recovery
A provider failed every consumer for a day and said so only in its journal
(hq issue 179). The provisioner loop now emits provisioner.failing after five
minutes without a success — create, check or secret — and repeats it every
fifteen; provisioner.recovered on the next success, on withdrawal, and on the
first success after a restart, so the controller can name it in status
(hq ADR 0224).
2026-10-06 00:13:42 +02:00
mesh-admin ed6384feb0 Merge pull request 'Remove resolv-conf (hq ADR 0223 part 2, step 2 of 2)' (#77) from retire-resolv-conf into main 2026-10-05 22:08:02 +00:00
jochen 60604fcd11 Give the uplink modules short slugs, so their identity fits a backend's limit
Requiring wildcard-resolution gave each a consumer identity; mesh_<machine>_networkmanager is
over the 20 characters a backend keeps, and the anchor's declaration, which carries every
consumer's grant, could not be composed.
2026-10-06 00:03:40 +02:00
mesh-admin ff61578e3e Merge pull request 'Give /etc/resolv.conf to the uplink's holder (hq ADR 0223 part 2, step 1 of 2)' (#76) from resolv-conf-to-uplink into main 2026-10-05 21:57:03 +00:00
jochen 138d9afd7b Rename hosts to hostname, which also writes /etc/hostname (hq ADR 0223)
Two files say one fact, the machine's name, and nothing owned /etc/hostname.
The name written is the operator's hostname setting, with no default: three
of four machines call themselves something other than their mesh name, and
renaming one is the operator's call. It takes effect at the next boot.
2026-10-05 23:43:14 +02:00
jochen dd124966ad Remove resolv-conf now that the uplink's holder writes resolv.conf (hq ADR 0223)
Merge only once resolv-conf is unassigned on every machine and forgotten.
2026-10-05 23:40:38 +02:00
jochen 73d6a51325 Give /etc/resolv.conf to the uplink's holder (hq ADR 0223)
The program that manages a machine's network is the one that would rewrite
the resolver file, so its module now writes it: networkmanager,
systemd-networkd and dhcpcd render the same template from the resolver's
holders. resolv-conf declares nothing for one release, so every machine
hands the file over in one apply; it is removed once unassigned everywhere.
2026-10-05 23:39:13 +02:00
mesh-admin 7b0b80ee05 Merge pull request 'postgres: port to Go, and install the extensions a consumer asks for (letta: vector)' (#75) from feat/postgres-go-extensions into main 2026-10-05 21:31:14 +00:00
jochen d8b4d20886 Port postgres to Go and install the extensions a consumer asks for
letta crash-loops on 'type "vector" does not exist': pgvector is not a
trusted extension, so only the provider's superuser can create it, and
the provisioner never did. A contribution may now name extensions; the
provider creates each (IF NOT EXISTS, available ones only) in the
consumer's database on every pass. Go per the standing rule for a
TypeScript module that changes. letta asks for vector.
2026-10-05 23:29:55 +02:00
jochen 7f491fd6bc Studio listens on every address, so its health check reaches it
Next.js binds the address HOSTNAME names; docker sets HOSTNAME to the container's id, so studio
answered only on its network address while its image's health check asks localhost, and it read
unhealthy while working. HOSTNAME=:: as the upstream compose file sets it.
2026-10-05 23:20:35 +02:00
jochen 737f42deb4 List every mesh resolver and no public one in resolv.conf
musl asks every nameserver at once and takes the first reply, so a public
resolver's NXDOMAIN for a mesh name beat the mesh's answer in every Alpine
container (hq ADR 0223). resolv-conf now renders /etc/resolv.conf from the
holders of mesh-dns-resolver, this machine first when it holds one; dnsmasq's
comments say the seat may have several holders.
2026-10-05 22:42:54 +02:00
jochen 964a4fdfbc docker: trust the mesh's registry from the runtime's own module (hq issue 190)
The controller's private network writes insecure-registries into daemon.json, a file this
module owns. The runtime's module states it instead, through ${seat:mesh-artifact-store:reach}
(hq ADR 0222), so the controller can stop generating its registry-trust resources.
2026-10-05 22:17:16 +02:00
jochen 4bf5eef2fd A machine's mesh name has no IPv6 address rather than no name (hq issue 262)
The resolver answered a machine's name only by wildcard, which says there is no such name when
asked for an IPv6 address; musl reads that as final, so Alpine containers could not find a
machine at all. A host record per machine answers that the name exists and has none, as the
hosts file the per-machine resolvers read used to.
2026-10-05 22:07:28 +02:00
jochen 875d2a0554 Retire resolved-split-dns: ADR 0196 chose no stub, and nothing assigns it
The resolv-conf comment pointed operators at it and at NetworkManager as
alternative claimants; it now says the uplink's holder is required beside it
(hq ADR 0220).
2026-10-05 21:57:04 +02:00
jochen 076455ec78 hosts in Go, with the machine's own name in its block (ADR 0199)
The tools were TypeScript; the mesh's modules are Go. The write to /etc/hosts is now staged and
moved into place rather than written over the live file.
2026-10-05 20:42:58 +02:00
jochen f20c4b749b The runtime's file is written by the runtime's module, not by what decides how a machine resolves (issue 190)
resolv-conf would have taken over dnsmasq's write into daemon.json — the same defect issue 190
names. docker, on every machine, now writes live-restore and reloads its own service. Log rotation
is left as each machine has it.
2026-10-05 20:42:58 +02:00
jschoubben 69d6b9066f The mesh's one resolver, what every node asks, and a node's hosts file (hq ADR 0194, 0196, 0199)
- dnsmasq holds mesh-dns-resolver: provides wildcard-resolution mesh-wide, forwards every declared
  zone (zones fact), listens on the private address and loopback only, reads no hosts file and no
  operator's files, and no longer writes the container runtime's dns.
- resolv-conf names the mesh's resolver by address, then 1.1.1.1, timeout 1, one attempt; it now
  holds the runtime's live-restore, which dnsmasq held and every node needs.
- resolved-split-dns routes the suffix to the mesh's resolver by address, not 127.0.0.1.
- hosts: new module holding node-hosts-file — the machine's own lines in its block of /etc/hosts,
  the operator's lines kept, changed by entries/add/remove through sudo -n.
2026-10-05 20:37:59 +02:00
jochen 0581256905 build-agent serves its seat's verbs: current, kill, pause, resume (hq ADR 0219)
The node-build-agent seat now promises these verbs, and a holder that does not
name them cannot hold it. The binary serving them is mesh-controller's
cmd/mesh-builder; merge only once a controller carrying the seat's new verbs runs,
since an older one refuses a claim naming verbs its row lacks.
2026-10-05 19:23:30 +02:00
jochen 428f5b8864 distribution: collect for real now that every kept archive is held
The dry run stood until the controller held each kept archive by a manifest
(mesh-controller #53) and an apply stopped reopening the collection window
(mesh-host #23, hq issue 224). The controller's collection command now
reports 134 of 134 kept archives held. hq issue 253.
2026-10-05 18:34:37 +02:00
jochen 52b81d524c gitea: a merge poll looks only at repositories that moved, one pass at a time
With the poll the only announcer of a merge, its cost showed: every 30 s it
asked every repository for its pull requests, a pass outlasted the tick, and
passes piled up beside each other — a merge was announced four and a half
minutes late, and two passes at once could each announce it. A pass now asks
only repositories updated since a minute before the last look, and the next
pass starts when this one ends. hq issue 250.
2026-10-05 18:12:11 +02:00
jochen c42f1ce45b systemd: port to Go, and read a system unit's journal as root
The journal verb ran journalctl as the operator account, which outside the
journal's group sees only its own entries: every system service read
'-- No entries --', and a person reached for a shell. The read now
escalates with sudo -n like the acts; ported to Go with every test. hq
issue 255.
2026-10-05 18:09:42 +02:00
mesh-admin 6a42bafb1c Merge pull request 'records: port to Go, and keep the checkout in a directory it owns (hq issue 251)' (#64) from feat/records-in-go into main 2026-10-05 15:55:58 +00:00
mesh-admin 053eba6950 Merge pull request 'gitea: announce a merge once, and read every page of its changed files (hq issues 250, 252)' (#63) from fix/a-merge-is-announced-once into main 2026-10-05 15:55:54 +00:00