A roster fact used to be a name from a closed list, each formatted in Go
here — node-names as a hosts file, node-zones as a resolver's zones. Every
new consumer (ssh's known_hosts, an authorized_keys) meant another formatter
in the control plane, in the consumer's own configuration language.
Now a fact is a path and a Go template over the roster view (this node, the
suffix, and every served name vs the machines). The mesh owns the data; the
module owns the format. /etc/hosts is a template on the network module;
dnsmasq's zones move to dnsmasq. The controller renders and reads neither.
WireGuard stays a computed generator: the overlay is the substrate delivery
rides on, and its config is topology, not a roster projection.
Output is byte-for-byte unchanged, pinned by the hosts golden tests and the
resolver tests that compose the real dnsmasq manifest.
Design 25 §6 says an enrolling node subscribes the inbox its own token derives.
It had none: `sub` was empty, so a node would publish its request and wait out
its timeout against a mesh that had answered — the handshake could not have
completed.
And there was one shared `enrolment` user, which cannot carry that inbox at all:
a permission belongs to a user, so an inbox per token means a user per token.
Named after the node, which **is** the token's id — a token is issued for a node
record, the mesh holds one live claim per record, and the node's name is the one
identifier both sides have before anything else is agreed. It is also exactly
what the other transport does, where the account is named after the node and the
secret is its password.
A nameless enrolment user is now refused rather than composed into
`_INBOX.enrol..>`: an empty subject token, and worse, one every nameless
enrolment user would share — which is one machine able to read the credentials
sealed to another.
Still to wire: something that composes one of these per live token. Nothing
composes enrolment users yet, on either bus — on the old one the account is made
imperatively through the broker's management API when a token is issued, and
here there is no management API, so issuing a token has to recompose the server's
configuration. That is the remaining half of enrolment on the new bus.
A host's account reaches no part of the JetStream API — correctly, because the
controller is the only writer of consumer definitions — so the object a node
reads its declarations through has to be waiting before the host binds to it,
and nothing created one. Named after the node, because the node's own ack grant
is `$JS.ACK.NODES.<node>.>` and a consumer named anything else is one the host
cannot acknowledge a delivery from.
No max-deliver, and a five-minute ack wait: a declaration is settled only after
the node has applied it and reported, which is minutes on a machine pulling
images, and the stream holds exactly one message per node — so there is nothing
to dead-letter, only one message to redeliver for as long as that node is away.
Asserted on start as well as created at enrolment, for the reason the streams
are: a mesh raised from a restored backup has node records and no consumers, and
a node whose consumer is missing hears nothing while everything else about it
looks correct.
The test sets the consumer against the grant the node actually gets, because
each of the three ways of getting it wrong is silent: a wrong name cannot ack, a
wrong filter reads another node's declarations, and a pull consumer is one a host
has no authority to bind.
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.
Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.
Three things the wiring forced into the open:
**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.
**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.
**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.
`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.
The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:
**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.
**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.
The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.
Design 25 §7. A person is not a module and holds no seat: nothing is
addressed to them, nothing is delivered to them, and they have no durable
consumer. What they have is permission to ask, as a list of tools or `*`
for an administrator.
Four properties the tests hold it to, each of which is a way of being
wrong that would not announce itself: a person reaches nothing but tools,
so one cannot claim a module said something; no ack subject, because
authority over a consumer that does not exist is authority nobody would
audit; no allow_responses, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two
people do not share an inbox.
Design 25 §2 builds enrolment around carrying the reply subject in the
payload, because a JetStream consumer claims the transport Reply field for
its own ack address. The whole handshake rests on it, so it is checked:
the caller asked for _INBOX.LCr3M83q... and the consumer saw
$JS.ACK.PROBE.probe_consumer... The design was right, and the workaround
is necessary rather than defensive.
Worth having as a test rather than a note: if a future server version
stopped doing this, enrolment would keep working and the reason for the
payload field would quietly become folklore.
Step 3.4. On the new bus there is no reply queue to declare and no
correlation to check: each account is granted one inbox prefix and no
other, so an answer cannot reach the wrong asker. That settles a cost
build.go records having paid — on a shared reply exchange every asker saw
every result, which is why the correlation was checked rather than assumed.
And a tool nobody serves says so at once rather than after the whole wait.
The difference between "that module is down" and "that tool is slow" is the
first thing a person asking wants, and both tests are against a real server
because both are claims about what the server does, not about this code.
RequestBuild stays as it is, and is a different shape on the new bus rather
than the same one: a build takes minutes, so it is work submitted to a
queue with the outcome returning to a reply subject the request carries —
the pattern design 25 §2 already sets for anything crossing a stream. It
touches the builder too, so it goes with that conversion.
introduces
Step 3.4, the consume side's hard part. The guarantee (ADR 0083) is that a
push the controller cannot record because its store is restarting is held
and retried — never dropped, never falsely acknowledged. Keeping the
delivery unacknowledged in memory becomes a nak with a delay: the server
holds it, the controller keeps no list of parked messages, and a controller
that restarts mid-window loses nothing it was holding.
That is a plain win and it introduces one problem. Holding in memory let
the controller drop an older report when a newer one for the same node
arrived, "because acting on it after the newer would undo the newer". A
naked message is the server's and comes back whatever happened meanwhile,
so the older report is redelivered after the newer was applied.
The answer was already in the message. A report carries `Declared`, the
digest of the declaration it is about, which exists because an earlier
attempt to order reports by time lost the race it invited. So supersession
stops being something the controller remembers and becomes something it
checks — the same shape as a node refusing a superseded declaration by
sequence (issue 107): ordering settled by what a message says, not by when
it arrived.
Pure, so the guarantee is testable without a bus, a store or a clock. Nine
tests, including that staleness is decided before the store is waited on —
a redelivery that lost its race must not hold a slot a current message
needs.
A declaration-level test composes the shipped networking module with a
resolver and asserts /etc/hosts arrives as mesh-wireguard.fact-node-names
with into: block and region-only content, that the resolver's restart-on
still names it, and that a resource's at passes through untouched — a
composition step dropping into would otherwise go unnoticed. plan --show
marks files written into, so a region is not read as the whole file.
The rollout order is spelled out: every node's host, the controller's
own included, must be block-aware before this controller ships (hq 128).
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.
What remains is the client library's own package name, which is its name.
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.
Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.
Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
/etc/hosts is the machine's: the distribution's localhost lines, the
operator's own entries, and marked blocks other tools maintain there.
Writing node-names whole replaced all of it the moment the private
network was taken, and every later write by those tools was lost at
the next machine joining. The node-names fact is now emitted with
into: "block", so the host owns only its marked region and keeps the
rest byte for byte. The region holds only the mesh's names: no header
claiming the file, no localhost, no 127.0.1.1 line — the floor was
never the mesh's to write. How a fact is written is a property of the
fact in the closed table; node-zones stays a whole file the mesh owns.
Sequencing: a host older than the block mode refuses the whole
declaration on an unknown into, so every host must be upgraded before
this controller is rolled out.
the-uplink joins the closed set as a node seat delivering nothing. Its
holder is the module for the machine's own network manager, and keeps
that manager from contradicting the mesh — the resolver file left to
resolv-conf, mesh0 left alone — without ever declaring a link. Held
per machine, so a machine running two managers is refused at
assignment rather than found by its resolver being rewritten. The
count test moves to fifteen; one test holds the seat's shape.
Every required header set, each value in the pinned shape, and the subject
derived the same way. Read from the sdk's conformance directory by sibling
path, never copied.
Design 29 §1's load-bearing rule: a module names its events, tools and
seats locally and the mesh derives the subject, so reorganising the subject
space leaves every manifest correct. It held by construction, and a rule
held by construction is one a later field breaks quietly.
novox/hq ADR 0118: the prefix is the reservation rule, so a module
declaring any mesh-* name is refused and there is no reserved-names list to
drift. Ten seats renamed in the table, the manifests that claim them, the
controller's own shipped manifests, and the tests.
Not the migration 0118 expected: a holding is derived at resolution from
manifests and never stored, so nothing recorded points at an old name. A
kept rename table tells a manifest written against one what it became —
kept rather than retired, because a module lives in its own repository and
may be registered long after the catalogue stopped using it.
**A seat is not the interface it delivers.** The git seat became mesh-git
and the git provision did not; likewise the package registry. A blanket
replace renamed both, and the failure read "the package registry is served
on <nil>", which does not say "you renamed an interface". A test now pins
every seat against what it delivers, and that neither name is also the
other.
Carries MTU from the reported tunnel (mesh-host#28) through inventory,
the overlay graph's TakeOver, into the generated config's [Interface].
A tuned path keeps its MTU across the takeover instead of regressing to
1420 and hanging transfers no ping would reveal. Two emit tests; a
tunnel with no MTU writes no line.
A home node behind NAT (no Endpoint → not Reachable) that took over a
tunnel must still listen on that tunnel's port: its LAN peers dial it
there. ListenPort was gated on Reachable, which conflated 'a peer dials
me here' with 'the hub can dial me' — so the takeover guard refused
overlay-up, and the guard's suggested remedy (re-place with an
endpoint) breaks a NAT'd node's path: it stops keepalive and hands the
hub a private LAN address to dial. TakeOver now carries the found
tunnel's port (already known to the controller), and the interface
listens on it when the node is not otherwise reachable. Two tests;
Endpoint-reachable nodes keep the old path unchanged.
Task 3.9's other half and 1.4's missing client. The derivation is pure and
unit-tested; only "does the server accept this" needs one running, behind
MESH_TEST_NATS so the ordinary suite stays offline.
A seat's work queue is created at registration, not assignment, so work
queues until a holder appears — a stream created at assignment would make
"the holder is not here yet" mean "your messages are gone". Named after the
seat, because the holder can change and the queued work must not care.
A holder's worker uses a queue group even though the seat guarantees one
holder: the seat is authority, the queue group is delivery, and tying them
together means the day somebody allows two holders every message is
processed twice with nothing reporting it.
One consumer per module carrying every filter, because its ack permission is
derived from its name.
And a real bug the live server caught: a durable name may not contain a dot,
but an ack subject is $JS.ACK.<stream>.<consumer>, so the single string that
read correctly inside the permission was rejected as a consumer name. Split
in two, beside the permission that has to match. Unfixed, the symptom would
have been every message redelivered forever with a permission list that
looks right — which is the failure design 25 §4 warns about.
The manifest carries seats and uses; registration refuses a mesh-* name, a
duplicate declarer, an undeclared uses or claim, a seat with no protocol, a
scope mismatch, and a holder that does not answer what its seat promises.
The parser stops judging unknown claim names, because it cannot: another
module may declare that seat, and one manifest cannot tell. The test that
encoded the old rule is rewritten to assert the refusal at registration, and
a new one pins the case the parser could not have distinguished.
Tools are declared for the first time, under their own key — serves already
means a provision's facts.
NATS refuses overlapping streams rather than double-storing, which is the
opposite of what the Overlaps comment claimed. The check still earns its
place — it names both streams at composition rather than one at apply — and
the refusal is what rules out a shared stream beside per-module ones.
Step 2.3 of novox/hq ADR 0116. The refusal is the resolver's existing one;
these pin it for this seat, including that a different bus implementation is
refused for the same reason — the property that lets the bus be replaced.
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).
Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.
Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
heartbeats live under that prefix and a stream of them competes for
retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
filter over a module's whole namespace would persist every tool call
Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.
Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.
NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.
Two things found by composing a real file rather than reading the design:
- a scoped inbox leaves a responder unable to reply, because the answer goes
to the caller's inbox. allow_responses is the answer — one reply to the
subject of a message actually received — and only principals that serve
are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
file's digest, so an order-dependent composer would reload the whole bus
on every controller restart. Covered by a test.
The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.
The tunnel the hub took over routes to machines the predecessor knows
by name and the mesh knew only by address — taking the resolver in that
state silences three machines at once. Now the operator states which
machine a carried address is (overlay name <address> <name>), the
statement rides tunnel_peer.named, and namesInTheMesh answers for named
not-yet-enrolled peers — one reading, so the hosts fact, a container's
hosts and the resolver cannot disagree. Enrolment verifies the word:
a machine enrolling under a named peer's key with a different name is
refused where the operator can read it, the stated name keeps the
carried address, and an enrolled peer's name is the node's — naming it
again refuses. The issue's rule holds: a name the predecessor answers
for keeps resolving until the machine behind it is a node.
The bus is the only broker (novox/hq ADR 0117): messaging is subjects on it,
scoped by a module's own emits/consumes, not a server handed out as a
provision. So the seat joins mesh-controller and the-catalogue in delivering
no interface. Full suite green.
One assignment of a module per node is now the rule, not a limitation —
the operator dropped the multi-assignment requirement, and the schema's
(node, module) key has been the decision since migration 0005. What
changed: Assign reports whether the assignment was new, and the command
says 'already runs — one node runs one of each (ADR 0115); nothing
changed' instead of printing 'is assigned' for a no-op, which read as
an action that happened. Idempotence stays: a repeat is exit 0, because
a script stating what is already true is not wrong.
checkResources compared paths as written, so ${dir:state}/server.env —
the same characters in every module, a different directory in each —
refused the first two placed modules that met. Paths are placed before
they are compared, under the default root, which keeps every real
collision: distinct modules' places are distinct under any one root,
and a module stating another's placed root is caught because a pathless
directory now owns its placed path in the comparison too.
Slice two of ADR 0112. A pathless directory saying place "." is the
assignment's one directory, <root>/<module> — to-be 27's shape — and
place never reaches the host, which parses strictly. The maps naming
where bindings, credentials and contributions land (binds, secrets,
own-secrets, receives, grants) fill against the placed directories at
composition, into fresh maps and a fresh module slice, because one
resolution composes for many nodes. The five absolute-path checks on
those maps accept a placed reference — resolution makes it absolute
before anything reads it — while certificate, operator-keeps and
accesses paths stay absolute-only: those are the operator's or another
vocabulary's. unknownDirRefs scans the maps too, and validates place
itself: only on a directory, only ".", never beside a stated path.
Found by the foundation tests validating the sibling catalogue: the
first conversion's blanket replace turned /var/lib/gitea/database.json
into ${dir:data}base.json — which resolves to the right path by pure
string concatenation. Production was saved by a coincidence; the
catalogue cleanup that follows spells it ${dir:state}/database.json.
The first executable slice of ADR 0112 / to-be 27, sized to what the
operator settled tonight: a module definition names no host path for
its own data. A directory resource may omit path; composition resolves
it to <root>/<module>/<id>, the root a node's setting on Rendering with
/var/lib as the default — which reproduces exactly the layout novox
converged to by hand. ${dir:<id>} names the place from a resource's
path, content, mounts, environment and env-files, the same shape as
${bound:…}. A directory that states a path keeps it and still answers
by name — that is the adopted-data placement, mssql its live case.
Resolved in the controller at composition, so the wire format and the
host change not at all; a reference naming no directory refuses at the
manifest and again at composition; nested fills are rebuilt, never
written into the manifest's own maps, because one manifest composes
for many nodes.
A private repository could not be built: the builder clones anonymously,
and had no way to say who it is. It already holds exactly one credential
to exactly the right place — the package-registry binding and its sealed
secret, one gitea user whose password answers npm and git alike — so a
clone now offers that, and nothing new is minted or carried.
Offered, never pushed: the credential is written as a git
credential-store file (0600, in the workspace, never argv) and named
with -c credential.helper, so git itself decides when it applies — only
on an authentication challenge, and only for the URL it was written
for, scheme, host and port included. A public repository clones exactly
as before; a repository on any other host is never shown it. The same
store rides along on an artifact's own context clone, so a private
module with a private context builds too.
Implements novox/hq ADR 0110 and 0111.
The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying
it delivers, and the record that made it one. A test asserts the count and a decision per entry, so
changing the set means finding the argument, as the host's vocabulary test does. The first set is
every seat already claimed — including the-private-network, which the network module claims from a
manifest composed in this repository's code, not from any module.json — plus npm-package-registry
(ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and
fails on any refused claim, so closing the set refuses nothing in use.
ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another
scope, and a delivering seat claimed by a module that does not provide what it delivers. A
malformed claim is refused once, for being malformed.
Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the
seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never
guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node.
A provider now carries the module it came from, because a provider is a (node, module) pair and the
pair is what tells a holder from a neighbour on the same machine.
The planner's second pass is now given the first pass's holdings. Without them, a node consuming a
seat-delivered provision was refused there, and a refused node's own claims dropped out of what the
mesh holds — letting a second holder of one of its seats pass unrefused.
`seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments
every time and never stored. Unheld seats are listed. A stored claim outside the set — possible
for a manifest registered before the set closed, since stored manifests are not re-validated — is
shown rather than hidden.
`build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is
composed at build time from the holder's node and what it serves for git; the recorded source is the
path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded.
Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An
address passed with --self is refused rather than recorded as a path.
Replaces three foundation tests that defended the builder's carried package binding. The catalogue
removed that binding when the builder began requiring the registry through a real grant, so the
tests were already failing on main; they now assert the builder requires what the npm seat delivers
and carries no copy of its own, and that the forge holds the npm and git seats.
Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new
inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on
main too.
route-proxy's own Dockerfile documents the shape it has always needed and
never had: 'the proxy source is not vendored here... the build context is
the mesh-controller repository root, and this Dockerfile compiles
./examples/route-proxy from it.' Nothing in the mesh could do that — the
build command clones one repository and builds every artifact from
within it, so route-proxy has never once been built through the pipeline,
consistent with it never having been assigned anywhere. Found attempting
exactly that build tonight: 'stat go.mod: file does not exist', because
the context was mesh-catalog, which does not have one.
An image artifact may now carry a context: {repository, ref}, cloned
fresh alongside the module's own tree. The recipe (Dockerfile) is still
read from the module's own directory, at the module's own commit — only
docker build's own context argument moves. Packaging and source stay
exactly as separate as route-proxy's own comment already said they were,
now for real.
TestTheForgesSshPortIsGivenByTheNumberTheForgeCallsIt exercised a settings
override from '2222' to 222 — but 2222 was never a real port anywhere,
just a mistake in gitea's own manifest (fixed alongside this: listens.port
is now 22, the container's real internal sshd port, matching every other
module's convention, and ports declares 222:22 directly — 222 has always
been the real, fixed public git-ssh port, needing no per-node override).
Split into two tests: the fixed default with no override, and a genuine
override case for a hypothetical node whose predecessor used a different
number, keyed correctly by 22.
TestTheResolverAndWhatAsksItComposeOnOneMachine set Rendering.Names but
FactNodeZones reads Rendering.Machines (novox/hq issue 111 split the two
apart: every name the mesh serves vs. the machines subset) — a loose end
from that merge, not exercised until now. Both are the same map in this
test's scenario, so both fields are set.
Every cutover done on novox tonight (drive, files, files-api, git,
keycloak, umami) dropped the <label>.<node>.internal alias HAL always
paired with the public hostname — found only when the operator tested it
by hand. Not a security boundary (a predecessor proxy served both as a
convenience, reaching a service over the VPN without a public TLS round
trip, not as access control), so restoring it is composing the same
convenience the same way the public name already is: <label> joined to
the node's own private address (r.At), independently of whether a public
domain exists to join the other half to.
composeName's signature changes (publicDomain, internalDomain) but its
shape does not — additive, label-gated, apex-aware, exactly mirroring the
public half it already did. A contribution the mesh writes both names
into is the entire fix; route-adapter and route-proxy pick up internal-
name whenever they're updated to serve it, not before, so this alone
changes nothing about what is live on any node yet.
ContributionsFrom settled to whichever of a module's several contributions to
one requirement sorted first, arbitrarily — the grant minted for it then
carried that contribution's label and port under a credential the OTHER
contribution's consumer never sees, and collided with that same
contribution's own entry from contributions() besides.
Confirmed live: minio's two route contributions (files-api, files) produced
three entries in route-adapter's received file — files-api twice, once
credentialed and once not, files not credentialed at all. Every
single-contribution module (gitea, keycloak, umami) already mints an unused
credential for `route` too — route never needs one, by its own
documentation — but with exactly one contribution to match there was nothing
to collide with, so it never surfaced.
Where a module contributes more than once, there is no single value to
settle on. The module still asks, still gets its one credential — a pair
credential is not a place for a label or a port anyway — and each named
contribution reaches the provider on its own, unchanged.
No cleanup needed for the secret already minted live for minio+route: the
sealed blob is a random pair credential unrelated to Values, which is
recomputed fresh on every plan/push regardless.
A module's contributes was map[string]map[string]any — one JSON object key
per requirement, structurally exactly one contribution to "route" ever.
minio needs two public hostnames (the S3 API and the console), which is
two different contributions to route from one module, and nothing let it
say so.
This is the same shape of problem ADR 0094 solved for secrets (a module
needing several values from one provider that gives one per pair):
contributes now accepts either the ordinary {label, port} object, or an
object of local names to several such objects. Detected per requirement
key by what's inside, since (unlike secrets' string-vs-object split) both
shapes are JSON objects: an ordinary contribution's fields are scalars, the
several-instance shape is local-name -> object. Confirmed against every
module.json in mesh-catalog before relying on that split.
Both route-proxy and the migration-era route-adapter already key generated
routers off the composed hostname (Values["name"]), not the module name,
so two contributions with the same From reach them as two independent
routes with no changes needed on the receiving side.
The map the control plane hands a resolution holds both: the machines, and every name
the mesh was told to route to whichever machine serves it. A container's hosts wants all
of it, so a routed name resolves to the proxy. A resolver's zones want only the machines:
told the mesh's suffix is its own it answers authoritatively for everything under it and
forwards none of it, so a routed name with the suffix appended — drive.example.test.internal
— is a name nobody will ever ask for, standing beside the machines and looking as real.
Found composing the resolver's first assignment on a live machine, before pushing it.
hq issue 111.
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.
A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.
The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.
The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.
Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:
- A predecessor spoke's tunnel names one peer, the hub, routed the whole
range; recording refused it and the whole enrolment failed. Range-routed
peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
so converging the hub would have renumbered the mesh and dropped the peers
still reaching it. They are facts of the tunnel record now, mode aside; the
takeover alone is declared to an adopted node. Converging the hub is refused
while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
with the tunnel, which would have the host stop the found interface and
raise the mesh's where no peer listens. The graph refuses to compose it,
naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
re-enrolling would rotate every key the mesh sealed credentials to. A node
now rekeys in a report, signed with its identity key over the key it
leaves, the key it takes and the tunnel; the mesh verifies against the live
key, refuses a stale or foreign proof, records key and tunnel, and moves a
hub to the tunnel's address. `overlay show` names the path for a hub that
found no tunnel.
Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
Beside each sealed connection genesis wrote, the port this machine put the
seat's holder at: `${seat:mesh-store:5432}` for the three stores,
`${seat:mesh-broker:…}` for the bus, the plain AMQP port and the management API.
Filled from the node's settings when the control plane composes its own
declaration; empty — the sealed value stands — when the mesh has nothing to add.
On its own, after the commit before it is built and running: a control plane
that does not know the placeholder passes it through as the value, and this
manifest is composed by whatever control plane is running when it is pushed.
The reader ignores an unfilled placeholder either way, and a test holds it to
ignoring exactly what this manifest says.
novox/hq 04-ISSUES/102
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.
The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.
A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.
novox/hq 04-ISSUES/102