Commit Graph
886 Commits
Author SHA1 Message Date
mesh-admin 6e4d12e5a4 Merge pull request 'Say which modules wait for a person's push; announce a merge's deleted files (hq ADR 0236)' (#95) from feat/core-upgrades-that-roll-back into main 2026-10-06 17:32:58 +00:00
jochen 1fd2914ae1 Say which modules wait for a person's push, and announce the files a merge deleted (hq ADR 0236)
With a gate on the first machine and a rollback after it, a module's build rolls out by
default. The ones kept back say why: the network path a rollback could not cross, the
providers every consumer on a machine drops with, and the stores holding the photos.
A merge's deleted files are announced, so a module whose manifest went is forgotten
rather than asked to build (the public-acme plan failure).
2026-10-06 18:39:13 +02:00
mesh-admin 7f99fb4a85 Merge pull request 'audit-logger, model-usage: retry a failed write and never lose the event (hq issue 276)' (#94) from fix/audit-and-usage-never-lose-an-event into main 2026-10-06 16:38:06 +00:00
jochen 08265a70ca audit-logger, model-usage: retry a failed write and never lose the event (hq issue 276)
Both caught a failed write and took the event, losing it silently; the SDK's rule is to throw when
the work was not done. A failed write now throws so the bus offers the event again, and is spooled
on disk at once; on its last delivery the spooled event is taken, and a background pass replays the
spool once writing works. The runtime does not pass the delivery count, so the spool counts failed
deliveries itself, across restarts. Over its bound (1000 events or 30 minutes) the last delivery is
no longer taken, so the bus gives it up and the controller raises max-deliveries - the one existing
condition that names a consumer which cannot keep up - while the spool still holds it.

Writes are idempotent by event: the trail skips an id it already wrote; the usage upsert keeps the
reading observed latest (migration 2), so a late replay never overwrites a newer one. Each module has
a status tool for the spool, declared as valuable data (ADR 0233). model-usage moves to the bundle
shape (ADR 0198) with its schema in a prepare step and numbered migrations; its old container shape
had no image. Both on mesh-sdk 0.1.13.

The log-only handlers of redis, mssql, mosquitto, mongodb, mesh-vault, showcase and the catalogue no
longer throw a TypeError on an event without a body.
2026-10-06 18:36:16 +02:00
mesh-admin f144f6eee8 Merge pull request 'Every module declares its data; the backup holder measures it (hq ADR 0233)' (#92) from feat/a-module-declares-the-data-it-holds into main 2026-10-06 15:05:51 +00:00
jochen 7524390cad Bound the holder's measuring; dump the database platform (hq ADR 0233)
Walks run at most daily and stop after ten minutes or two million files; datasets are read from
their counters and large items from their top level only. The database platform's tables are
dumped with pg_dumpall rather than copied as live files.
2026-10-06 17:00:23 +02:00
jochen 685cb1cb1b Declare every module's data; the backup holder measures it (hq ADR 0233)
Backup lines are derived from each module's data section instead of written by hand; the holder
measures declared items, reads the array under them, and deletes a retired item only after a last
restore point; the Go providers say each held consumer's size so an empty replacement is seen.
2026-10-06 16:47:49 +02:00
mesh-admin 8f5a75ab9b Merge pull request 'Fold public-acme into route-proxy; drop dhcpcd and cloudflare-dns (hq ADR 0226)' (#84) from feat/route-proxy-names-its-public-issuer into main 2026-10-06 13:04:29 +00:00
jochen 57d524f4ef Fold public-acme into route-proxy; drop dhcpcd and cloudflare-dns (hq ADR 0226)
public-acme ran nothing and had one consumer. The proxy now states the issuer itself, byte for byte
what the binding rendered, so its account directory and every certificate stay put. dhcpcd and
cloudflare-dns are assigned nowhere and nothing requires what they provide.
2026-10-06 14:59:43 +02:00
mesh-admin 368fa2f45e Merge pull request 'postgres, keycloak: retire a consumer, delete only on a person's word (hq ADR 0230)' (#91) from feat/retired-consumers into main 2026-10-06 12:51:35 +00:00
jochen 2c15074a0b Retire only once the same answer has held ten minutes as well as five passes
Five passes are twenty-five seconds, shorter than a controller restart, a store
reconnecting or a file half written; the operator asked for both (hq ADR 0230).
2026-10-06 14:28:21 +02:00
jochen e68ef88333 Retire a consumer the mesh stops asking for, and delete only on a person's word (hq ADR 0230)
The hourly release of ADR 0229's brake still ended in the mesh acting alone on
a mistake. A consumer now stays active until the same unasked set holds for
five passes, waits for a person past three or half of those held, is disabled
and marked rather than withdrawn, comes back as it was when asked again, and is
deleted only through the provider's delete tool. The backend keeps the mark, so
a restart forgets nothing and finds what was withdrawn before.
2026-10-06 13:54:10 +02:00
mesh-admin 4554e18279 Merge pull request 'Brake a withdrawal larger than its bound (hq to-be 45 Phase 2, ADR 0227 rule 4)' (#90) from feat/a-core-that-cannot-fail-silently-phase-2 into main 2026-10-06 10:43:11 +00:00
jochen 4238dd8616 Name the SDK harness the Go loop follows: 0.1.11, with the withdrawal brake 2026-10-06 12:29:21 +02:00
jochen 2c151e7c21 Brake a withdrawal larger than its bound (hq to-be 45 Phase 2, ADR 0227 rule 4)
Issue 241 withdrew seven consumers in one pass on one misread file; its fix
refuses a file it cannot read, and a file read whole that names nobody
still withdraws everybody. A pass that would withdraw more than one
consumer at once, or more than half of those it holds, now withdraws
nothing: each consumer it kept is announced provisioner.failing with the
class withdrawal-braked, so the controller raises it as a condition and
the operator is told, and one is let go each hour while the mesh goes on
not asking for them. A consumer asked for again is kept and said recovered.
Postgres and keycloak carry the harness identically.
2026-10-06 12:29:21 +02:00
mesh-admin 12569ed085 Merge pull request 'Mark the OpenAI keys letta and supabase hold as issued outside the mesh' (#89) from feat/a-given-secret-lives-until-the-first-good-start into main 2026-10-06 10:27:21 +00:00
jochen b1bd5c861e Mark the OpenAI keys letta and supabase hold as issued outside the mesh
The controller now replaces a given at-start secret after the module's
first good start and on rotate (hq ADR 0228); a key only OpenAI can
issue must say so, or a fresh random value would take its place.
2026-10-06 12:13:57 +02:00
mesh-admin 8c2c9ace0e Merge pull request 'messenger: history is never news — read the open conditions first, coalesce bursts (hq issue 271)' (#88) from fix/messenger-history-is-never-news into main 2026-10-06 10:06:27 +00:00
jochen a4b92c1306 messenger: never say history as news — read the open conditions first, coalesce bursts
A consumer made on 2026-10-06 was handed three hours of raised-and-cleared
conditions at once and the holder said each as new: 20 desktop notifications
in a second. What is said is now decided by the controller's open set and by
an event's own time, never by its arrival; bursts are one message, the
desktop gets warnings at most every 15 min, the cap is said once, and the
first minute after start says only the urgent conditions still open.
Replays of that morning's 96 events are tests. Also drops two committed
binaries. (novox/hq issue 271)
2026-10-06 12:03:19 +02:00
mesh-admin aabc8aa039 Merge pull request 'supabase: make logflare's stored key and backends follow its environment' (#87) from fix/supabase-logflare-reconcile into main 2026-10-06 09:50:28 +00:00
jochen 908d45864a supabase: make logflare's stored key and backends follow its environment
logflare 1.4.0 copies LOGFLARE_API_KEY into its default user and
POSTGRES_BACKEND_URL into every source's backend once, when it creates
them, and never reads either again. On ace every source still points at
db:5432, which resolves nowhere, so analytics stored no logs; and the key
that leaked into the log before #85 could not be replaced, because the
mesh had to mark it "applied".

The start script now runs logflare's migrations and then reconcile.exs
through `logflare eval`, before logflare starts: it uses logflare's own
Users and Backends contexts to set the default user's key to
LOGFLARE_API_KEY (clearing old_api_key) and to point each postgres
source backend at POSTGRES_BACKEND_URL. It changes nothing that already
matches, says what it changed without printing the key or a password,
and stops the start when it cannot finish.

With the key taken at every start, the secret is "at-start" and `secret
rotate` works. analytics restarts on its env file and studio on its env
file too, so a rotated key reaches every reader.
2026-10-06 11:48:57 +02:00
mesh-admin c12d364a68 Merge pull request 'The operator-channel and the watcher's watcher (hq to-be 45 phase 1)' (#86) from feat/the-mesh-says-when-it-is-wrong into main 2026-10-06 08:37:48 +00:00
jochen 0ae7933d54 Tell the operator what the mesh finds wrong, and watch the watcher (hq to-be 45 phase 1)
The mesh noticed 48 core failures in six days and told nobody (ADR 0227).
messenger holds the operator-channel seat: it consumes the controller's
condition events and sends them to Telegram and the desktop notifier,
deduplicated by key, reminded once, edited on clear, capped at 20 an hour
with the rest folded, and refusing anything carrying an address, a path or
a secret. mesh-watcher, on a machine other than the control node, sends to
Telegram directly when the self-check heartbeat or the bus goes silent.
2026-10-06 09:35:55 +02:00
mesh-admin 495bb88111 Merge pull request 'supabase: keep logflare's API key out of its log (hq issue 268)' (#85) from fix/supabase-logflare-key-out-of-logs into main 2026-10-06 00:34:13 +00:00
jochen b87dc29706 supabase: keep logflare's API key out of its log (hq issue 268)
vector handed logflare its API key as ?api_key= in every sink URL, and
logflare 1.4.0 prints a failed request's whole URL in its Plug.Cowboy
error report. Its ingest fails on every request here, so the key was in
the analytics log about every ten seconds. The report is an error, so no
log level hides it.

Every sink now sends the key in the x-api-key header, which logflare
reads first. A start script refuses to start logflare while the vector
config it is given still puts the key in a URL.

The key is marked "applied", not "at-start": logflare writes it into its
default user once and never updates it, so the mesh must not rotate it
by restarting.
2026-10-06 02:33:32 +02:00
mesh-admin e5cb7b071c Merge pull request 'State each provision's identity bound (hq issue 263, ADR 0225)' (#83) from fix/263-identity-bounds-per-provision into main 2026-10-06 00:29:56 +00:00
mesh-admin 7fad19a764 Merge pull request 'letta: keep its passwords out of its log; docker: find and hide secrets a container printed (hq issue 268)' (#82) from fix/letta-secrets-out-of-logs into main 2026-10-06 00:25:32 +00:00
jochen 4769dadf63 State each provision's identity bound (hq issue 263)
Consumers were held to an S3 access key's 20 characters whatever they
required. Each provider now says what its backend keeps: minio 20,
PostgreSQL and MongoDB and DNS 63, Gitea 40, a mailbox 64, SQL Server 128,
Keycloak 255, unbounded where the store has no limit, and none for the
resolver and route provisions, which keep no name of their consumers.
Needs the controller that reads the field (mesh-controller, ADR 0225).
2026-10-06 02:16:27 +02:00
jochen bbb67e41a0 docker: find and hide secrets a container printed into its log (hq issue 268)
letta printed two passwords into its log for weeks and nothing noticed,
and docker_logs handed them to whoever asked. docker_secrets_in_logs
compares each container's recent lines with the secret-named values of
its environment, the passwords in its URIs, and any URI carrying a
password, and names what it found by container, module and variable -
never the value. docker_logs redacts the same values before answering.
2026-10-06 02:13:42 +02:00
jochen f9f27d4878 letta: keep its database and server passwords out of its log (hq issue 268)
letta 0.6.8 prints LETTA_PG_URI whole (startup.sh, alembic, server.py)
and its server password when it starts in secure mode, so both were in
the container's log on every one of its restarts. Newer letta still
prints both, and neither is a log level.

The URI now names no password: libpq reads it from a mounted pgpass
file (PGPASSFILE). The one print of the server password is rewritten by
a start script before the server starts, and the script refuses to start
letta if that print, or a password in the URI, is still there - a letta
that does not start says why; one that leaks says nothing.

Both own secrets say they are read at start, so `rotate` can replace the
server password the mesh made.
2026-10-06 02:13:42 +02:00
mesh-admin cc2f19123a Merge pull request 'nats: run 2.11.17, which hands a consumer with several filters every message (hq issue 266)' (#81) from fix/nats-multi-filter-consumers-skip into main 2026-10-05 23:40:11 +00:00
jochen 0275c2eeac nats: run 2.11.17, which hands a consumer with several filters every message (hq issue 266)
On 2.10.29 such a consumer was moved past a message now and then without
handing it over; the controller's events consumer has seven filters, and a
merge on the stream never reached it. The test reproduces the skip on 2.10.29
and keeps the image's release equal to the server it tests.
2026-10-06 01:29:20 +02:00
mesh-admin 78328d4ab2 Merge pull request 'keycloak: port to Go and repair a refused admin; providers announce a failing consumer' (#80) from feat/identity-provider-admin-safety-nets into main 2026-10-05 22:39:02 +00:00
jochen 48d4188927 keycloak repair: remove the temporary admin even when the bootstrap failed after making it
The bootstrap once created the temporary admin and then failed on a held port; marked only after
it succeeded, the cleanup did not know the admin existed and left it.
2026-10-06 00:17:35 +02:00
mesh-admin 5c2157b81c Merge pull request 'Rename hosts to hostname, which also writes /etc/hostname (hq ADR 0223 part 3)' (#78) from hostname-module into main 2026-10-05 22:15:29 +00:00
jochen 77fb1ecfb2 keycloak: port to Go and repair an admin that refuses the mesh's secret
Twice the identity provider's admin kept an older password than the one the
mesh minted (an adopted, then a moved database), and the provisioner failed
every consumer until it was repaired by hand (hq issue 179). The module now
checks the admin's login and repairs a refusal itself through the server's
bootstrap command, verifies, brakes a failed repair and announces it, and
stops asking the server while refused. Ported to Go to change it.
2026-10-06 00:13:42 +02:00
jochen 6f1e2f5a0d postgres: announce a consumer failed for minutes, and its recovery
A provider failed every consumer for a day and said so only in its journal
(hq issue 179). The provisioner loop now emits provisioner.failing after five
minutes without a success — create, check or secret — and repeats it every
fifteen; provisioner.recovered on the next success, on withdrawal, and on the
first success after a restart, so the controller can name it in status
(hq ADR 0224).
2026-10-06 00:13:42 +02:00
mesh-admin ed6384feb0 Merge pull request 'Remove resolv-conf (hq ADR 0223 part 2, step 2 of 2)' (#77) from retire-resolv-conf into main 2026-10-05 22:08:02 +00:00
mesh-admin 0e072b05c0 Merge pull request 'Uplink modules get short slugs (nm, networkd)' (#79) from fix/uplink-modules-have-short-slugs into main 2026-10-05 22:03:45 +00:00
jochen 60604fcd11 Give the uplink modules short slugs, so their identity fits a backend's limit
Requiring wildcard-resolution gave each a consumer identity; mesh_<machine>_networkmanager is
over the 20 characters a backend keeps, and the anchor's declaration, which carries every
consumer's grant, could not be composed.
2026-10-06 00:03:40 +02:00
mesh-admin ff61578e3e Merge pull request 'Give /etc/resolv.conf to the uplink's holder (hq ADR 0223 part 2, step 1 of 2)' (#76) from resolv-conf-to-uplink into main 2026-10-05 21:57:03 +00:00
jochen 138d9afd7b Rename hosts to hostname, which also writes /etc/hostname (hq ADR 0223)
Two files say one fact, the machine's name, and nothing owned /etc/hostname.
The name written is the operator's hostname setting, with no default: three
of four machines call themselves something other than their mesh name, and
renaming one is the operator's call. It takes effect at the next boot.
2026-10-05 23:43:14 +02:00
jochen dd124966ad Remove resolv-conf now that the uplink's holder writes resolv.conf (hq ADR 0223)
Merge only once resolv-conf is unassigned on every machine and forgotten.
2026-10-05 23:40:38 +02:00
jochen 73d6a51325 Give /etc/resolv.conf to the uplink's holder (hq ADR 0223)
The program that manages a machine's network is the one that would rewrite
the resolver file, so its module now writes it: networkmanager,
systemd-networkd and dhcpcd render the same template from the resolver's
holders. resolv-conf declares nothing for one release, so every machine
hands the file over in one apply; it is removed once unassigned everywhere.
2026-10-05 23:39:13 +02:00
mesh-admin 7b0b80ee05 Merge pull request 'postgres: port to Go, and install the extensions a consumer asks for (letta: vector)' (#75) from feat/postgres-go-extensions into main 2026-10-05 21:31:14 +00:00
jochen d8b4d20886 Port postgres to Go and install the extensions a consumer asks for
letta crash-loops on 'type "vector" does not exist': pgvector is not a
trusted extension, so only the provider's superuser can create it, and
the provisioner never did. A contribution may now name extensions; the
provider creates each (IF NOT EXISTS, available ones only) in the
consumer's database on every pass. Go per the standing rule for a
TypeScript module that changes. letta asks for vector.
2026-10-05 23:29:55 +02:00
mesh-admin 0515db043a Merge pull request 'supabase: studio listens on every address, so its health check reaches it' (#74) from fix/studio-listens-where-its-health-check-asks into main 2026-10-05 21:20:40 +00:00
jochen 7f491fd6bc Studio listens on every address, so its health check reaches it
Next.js binds the address HOSTNAME names; docker sets HOSTNAME to the container's id, so studio
answered only on its network address while its image's health check asks localhost, and it read
unhealthy while working. HOSTNAME=:: as the upstream compose file sets it.
2026-10-05 23:20:35 +02:00
mesh-admin b4b86c1452 Merge pull request 'resolv-conf lists every mesh resolver and no public one (hq ADR 0223)' (#73) from feat/the-mesh-has-two-resolvers into main 2026-10-05 20:50:53 +00:00
jochen 737f42deb4 List every mesh resolver and no public one in resolv.conf
musl asks every nameserver at once and takes the first reply, so a public
resolver's NXDOMAIN for a mesh name beat the mesh's answer in every Alpine
container (hq ADR 0223). resolv-conf now renders /etc/resolv.conf from the
holders of mesh-dns-resolver, this machine first when it holds one; dnsmasq's
comments say the seat may have several holders.
2026-10-05 22:42:54 +02:00