Step 1.1 and 1.2 of novox/hq ADR 0116. The server is a built artifact rather than the upstream image directly, because it needs an entrypoint of its own: the host can only recreate a container, and recreating the bus for every permission change drops every connection and every in-flight ack. nats-server reloads on SIGHUP by itself, so the config is mounted as a directory (not digest-tracked, hq issue 103) and the entrypoint watches the one file. Verified against the real server, not assumed: a user added to the config connects, a revoked one is refused, both within one poll interval, with the container's PID and restart count unchanged and "Reloaded: accounts" in its log. Two corrections found by checking rather than reading: - the seat delivers nothing now (hq ADR 0117), and the controller's parser refused the manifest until it did — "nats claims mesh-broker, whose holder answers for amqp, and nats does not provide amqp" - pinned to the multi-arch index digest; the first pin was the amd64 manifest, which builds here and fails on any other architecture
62 lines
2.9 KiB
Bash
62 lines
2.9 KiB
Bash
#!/bin/sh
|
|
# nats's entrypoint: run the server, and reload it in place when the mesh rewrites its
|
|
# configuration.
|
|
#
|
|
# **Why this exists inside the module** (novox/hq design 25 §5). The controller composes every
|
|
# account and permission into one configuration file, and that file changes whenever a module is
|
|
# added, reassigned, or a person's access is granted or revoked — which is often, and on the one
|
|
# server everything else depends on. The host has no way to say "reload this container": a
|
|
# container resource has `restart-on` and nothing else, and a container's `restart-on` means
|
|
# *recreate* — every connection dropped and every in-flight JetStream ack lost, mid-flight, for a
|
|
# permission change. `reload-on` is real but it is a *service* field, not a container's.
|
|
#
|
|
# nats-server already reloads its own configuration on SIGHUP — accounts, permissions, everything
|
|
# the mesh composes — without dropping a connection. That is the server's own documented
|
|
# capability, not something built for the mesh. So the configuration is mounted as a directory
|
|
# (a directory's contents are not digest-tracked the way a directly-mounted file's are, novox/hq
|
|
# issue 103), and this watches the one file inside it and signals the server itself. The host's
|
|
# only job is what it already does for any directory: keep the file's content current. Nothing
|
|
# here is declared `restart-on` or `reload-on`.
|
|
set -eu
|
|
|
|
CONF="${MESH_NATS_CONF:-/etc/nats/nats.conf}"
|
|
POLL="${MESH_NATS_CONF_POLL_SECONDS:-5}"
|
|
|
|
# The controller writes the configuration as part of the same declaration that creates this
|
|
# container, but the two are not ordered against each other. Waiting is correct and starting
|
|
# without one is not: nats-server would come up with its compiled-in defaults — no TLS, no
|
|
# accounts, every subject open to anyone who can reach the port — and then be reloaded into
|
|
# correctness a moment later. A bus that is briefly open to everything is not a bus that is
|
|
# briefly wrong; it is an open bus.
|
|
while [ ! -s "$CONF" ]; do
|
|
echo "[nats] waiting for the mesh to compose $CONF"
|
|
sleep 1
|
|
done
|
|
|
|
digest() { sha256sum "$CONF" 2>/dev/null | cut -d' ' -f1; }
|
|
|
|
nats-server --config "$CONF" "$@" &
|
|
server=$!
|
|
|
|
# Forward a stop to the server and let it drain, rather than dying and leaving it orphaned as
|
|
# PID 1's child.
|
|
stop() { kill -TERM "$server" 2>/dev/null || true; }
|
|
trap stop TERM INT
|
|
|
|
last=$(digest)
|
|
while kill -0 "$server" 2>/dev/null; do
|
|
sleep "$POLL"
|
|
now=$(digest)
|
|
# An empty digest means the file is mid-write or briefly gone. Reloading on that would hand the
|
|
# server a truncated configuration; the next tick sees the finished one.
|
|
[ -n "$now" ] || continue
|
|
if [ "$now" != "$last" ]; then
|
|
last=$now
|
|
echo "[nats] configuration changed; reloading in place"
|
|
kill -HUP "$server" || true
|
|
fi
|
|
done
|
|
|
|
# `wait` on an already-exited child still yields its status, which becomes this container's.
|
|
wait "$server"
|