Commit Graph
482 Commits
Author SHA1 Message Date
jschoubben f4bcb320fe contributes: a module may answer one requirement several times
A module's contributes was map[string]map[string]any — one JSON object key
per requirement, structurally exactly one contribution to "route" ever.
minio needs two public hostnames (the S3 API and the console), which is
two different contributions to route from one module, and nothing let it
say so.

This is the same shape of problem ADR 0094 solved for secrets (a module
needing several values from one provider that gives one per pair):
contributes now accepts either the ordinary {label, port} object, or an
object of local names to several such objects. Detected per requirement
key by what's inside, since (unlike secrets' string-vs-object split) both
shapes are JSON objects: an ordinary contribution's fields are scalars, the
several-instance shape is local-name -> object. Confirmed against every
module.json in mesh-catalog before relying on that split.

Both route-proxy and the migration-era route-adapter already key generated
routers off the composed hostname (Values["name"]), not the module name,
so two contributions with the same From reach them as two independent
routes with no changes needed on the receiving side.
2026-09-24 18:36:08 +02:00
jschoubben caf9746759 mesh-controller: mount the broker TLS directory bind, not the old named volume
Found checking whether the named volumes mesh-catalog PR #54/#55 replaced
are actually unused before considering them safe to remove -- this repo
has its own independent volumes declaration for the same TLS material
(mesh-controller reads it directly, not through lavinmq's own resource),
and it still named the old mesh-broker-tls volume.

Right now the content is identical -- copied once during the conversion.
If the cert ever rotates, lavinmq writes the new directory and this would
keep reading stale content from the volume nothing else updates.

Checked both repos for any other reference to the four converted volume
names (mesh-store-data, mesh-broker-data, mesh-broker-tls,
mesh-registry-data): this was the only one.
2026-09-24 17:19:48 +02:00
jschoubben 6090953843 Merge pull request 'Tell the resolver the machines, not the names the mesh merely serves' (#53) from fix/the-resolver-is-told-machines-not-routes into main 2026-09-23 23:32:22 +00:00
jschoubben 2277583e99 Tell the resolver the machines, not the names the mesh merely serves
The map the control plane hands a resolution holds both: the machines, and every name
the mesh was told to route to whichever machine serves it. A container's hosts wants all
of it, so a routed name resolves to the proxy. A resolver's zones want only the machines:
told the mesh's suffix is its own it answers authoritatively for everything under it and
forwards none of it, so a routed name with the suffix appended — drive.example.test.internal
— is a name nobody will ever ask for, standing beside the machines and looking as real.

Found composing the resolver's first assignment on a live machine, before pushing it.
hq issue 111.
2026-09-24 01:31:11 +02:00
jschoubben 6bf42025e1 Merge pull request 'Give the resolver the mesh's suffix as a local domain and a module its machine's address' (#52) from convert/dnsmasq-from-hal into main 2026-09-23 23:13:03 +00:00
jschoubben 0d8264ff55 Give the resolver the mesh's suffix as a local domain and a module its machine's address
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.

A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.

The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.

The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.

Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
2026-09-24 01:10:15 +02:00
jschoubben 6073e94a4f Merge pull request 'Adopt the predecessor's tunnel in place: its range, its address, its peers (hq ADR 0105)' (#49) from feat/adopt-the-tunnel into main 2026-09-23 22:38:31 +00:00
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00
jschoubben 7ef7669c0c Merge pull request 'An address is read from the node's settings where it is used, never recorded with a port (hq issue 102)' (#50) from fix/addresses-follow-the-node into main 2026-09-23 21:55:03 +00:00
jschoubben cdd3638312 The control plane's own manifest says where the node put the store and the broker
Beside each sealed connection genesis wrote, the port this machine put the
seat's holder at: `${seat:mesh-store:5432}` for the three stores,
`${seat:mesh-broker:…}` for the bus, the plain AMQP port and the management API.
Filled from the node's settings when the control plane composes its own
declaration; empty — the sealed value stands — when the mesh has nothing to add.

On its own, after the commit before it is built and running: a control plane
that does not know the placeholder passes it through as the value, and this
manifest is composed by whatever control plane is running when it is pushed.
The reader ignores an unfilled placeholder either way, and a test holds it to
ignoring exactly what this manifest says.

novox/hq 04-ISSUES/102
2026-09-23 23:49:55 +02:00
jschoubben e07b56ce43 An address is read from the node's settings where it is used, never recorded with a port
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.

The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.

A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.

novox/hq 04-ISSUES/102
2026-09-23 23:49:31 +02:00
jschoubben 1b5ccf4165 Merge pull request 'The artifact store's seat is one per mesh, and the test says so from the catalogue' (#51) from fix/the-artifact-store-is-one-per-mesh into main 2026-09-23 21:42:33 +00:00
jschoubben 26690d89f1 The artifact store's seat is one per mesh, and the test says so from the catalogue
Review of the registry work found the seat node-scoped: a second `distribution` on another
machine resolved cleanly there, and only afterwards did the mesh notice `artifact-store`
offered by two nodes, with every consumer elsewhere refusing to choose. A node-scoped
requirement with one candidate installs that candidate, so anything that wanted the store
beside it would have raised a fresh, empty store on the wrong machine first.

The claim is mesh-scoped in mesh-catalog now; this holds the catalogue's manifest to it —
a second store anywhere is refused by name, where it is assigned.
2026-09-23 23:40:53 +02:00
jschoubben 3c836f0abb Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.

The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.

Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
2026-09-23 23:26:34 +02:00
jschoubben 8fb32d7ee0 Merge pull request 'A node may move a port a module publishes as a mapping's machine side' (#47) from fix/move-a-published-machine-port into main 2026-09-23 00:33:06 +00:00
jschoubben 7d6f37af54 One entry per mapping, under the name the module itself uses
Review found the first pass aliased its answer under both ends of a mapping, which is
wrong wherever two mappings share a number: the alias lands on a key belonging to another
mapping, the later write wins, and the filter and the container then disagree — the very
fault this change exists to close. Two reproduced cases: a module publishing 8080:80 beside
9090:8080 had an explicit setting silently overwritten; a module publishing 4001:80 beside
4002:80 composed both containers onto one machine port, where before it was safely refused.

Now a mapping's answer is filed once, under the end the module names in its listens — the
number the plan, the filter, the openings, the guard and the consumer all ask for — and a
key that names two mappings is refused in the same words as a setting that does.

Also: the guard assertion in the end-to-end test failed open when the resource was absent;
the plan-mirroring helper now says it stands in only where the plan does not allocate, and
the assertions it feeds are narrowed to the port under test.
2026-09-23 02:32:28 +02:00
jschoubben 58644fd282 A node may move a port a module publishes as a mapping's machine side
A module publishing `2222:22` — the machine's own ssh daemon holds 22, so
the module takes 2222 and says so in `listens` — could not be moved. The
setting was read against the last segment of each mapping alone, so the
number the module uses everywhere else was refused as a port it does not
publish, and the node's every push failed for as long as the setting was
stored. The one key that was accepted, the container's own port, was then
read only when the container's mapping was rewritten: the mapping moved
and the ports map, the filter, the adopted node's openings, its guard and
what a consumer is told all stayed on the number the software had left.

Either end of a mapping now names it, and a given port comes back under
both, so every reader finds the same number under the key it holds.
Ambiguity is refused where it is real — one number naming two different
mappings, or the two ends of one mapping given two different numbers.
2026-09-23 02:08:35 +02:00
jschoubben 91a41b7d20 Merge pull request 'A container's environment follows a moved port, as a file already does (hq issue 088)' (#46) from feat/forge-address into main 2026-09-22 23:30:06 +02:00
jschoubben f0049190d7 Prove the untouched environment is the same map, not an equal one 2026-09-22 23:29:28 +02:00
jschoubben 7352c846dd A module is told its port in a container's environment too (hq issue 088)
${port:…} answered only inside a file's content, and the one place a module
routinely writes its own address is a container's `env` — where the literal is
wrong on every node whose assignment differs from the manifest's number, and
wrong again on a node given that port as a setting (ADR 0100). Nothing checked
it: the value is a string like any other, and it fails at runtime, on one node.

Filled by the control plane, like a bound value: a port is not secret, so there
is nothing for the host to be the only witness of and it learns no new field.
That is the line ADR 0086 draws — its objection is to a secret being in an
environment at all, not to who fills one in — so a port crosses it and a
credential still does not. Same guard as before: a port the module never said it
listens on is refused, now naming the container and the variable.

The env map is the catalogue's, shared by every node running the module, and the
resource around it is a shallow copy, so a filled value goes into a fresh map —
otherwise the first node composed writes its own port into the manifest and
every node after it is told that one.

Inert on the catalogue as it stands: ${port:…} is written in one other place in
it, a file. Renamed off _files, which this no longer is.
2026-09-22 22:36:28 +02:00
jschoubben 45425702d5 Merge pull request 'Hold the forge to being an ordinary provider, in tests (hq issue 085)' (#45) from feat/packages-port into main 2026-09-22 21:59:52 +02:00
jschoubben 729537e745 Tie the builder's carried binding to what the forge serves, and hold at shut
The builder carries a binding because at genesis nothing provides
`package-registry` to resolve one from; once the forge is a module the same
consumer is told what the forge serves. Nothing held the two to the same number,
so the catalogue could drift into dialling one port before the forge is assigned
and another after.

And `at` is now protected, for the reason it had to be: a setting that moves it
points the builder, and the registry password it sends, at a host somebody else
chose.

novox/hq 04-ISSUES/085
2026-09-22 21:56:51 +02:00
jschoubben 0e413e3e7a Hold the forge's port to the same rule as every other provider's
The package registry was the one foundation port not resolved from what its
module serves. Nothing in the controller had to change for it — `ports` on the
forge moves its container, what it serves and what consumers are told, and the
builder's carried binding is settable like any other mergeable file — but
nothing said so, which is how it came to be special in the first place.

Two tests over the catalogue's own manifests: the forge's port is given on a
node and reaches what it serves, and the builder's carried binding takes the
port from the node while keeping who the binding is with.

novox/hq 04-ISSUES/085
2026-09-22 21:40:08 +02:00
jschoubben 252186d90c Merge pull request 'Adoption mode: a node in use is adopted before it is converged (hq ADR 0100–0103)' (#44) from feat/adoption-mode into main 2026-09-22 21:01:51 +02:00
jschoubben fad8b30e43 Say that the guard names the runtime's bridges where the filter names their addresses (hq ADR 0103) 2026-09-22 19:51:10 +02:00
jschoubben 1096299e06 Hold a converged declaration to the bytes main sends, captured from it, rather than to re-marshalling itself (hq ADR 0100) 2026-09-22 19:51:10 +02:00
jschoubben 379f459498 Hold the filter module to reloading its rules and restarting only on its units (hq ADR 0102) 2026-09-22 19:47:43 +02:00
jschoubben d05a5e87af Hold the node while an assignment is recorded, so none lands between a preview and the flip (hq ADR 0100) 2026-09-22 19:47:23 +02:00
jschoubben 01814854b8 Refuse the flip on an account naming nothing reachable, and mark such an account partial in the preview (hq ADR 0100) 2026-09-22 19:47:23 +02:00
jschoubben 2cf8739a84 Clear the whole account an adopted node gave when it converges (hq ADR 0100) 2026-09-22 19:45:35 +02:00
jschoubben 87ecc9326e Refuse a port given for the whole mesh where it is set, not at every node's composition (hq ADR 0100) 2026-09-22 19:45:35 +02:00
jschoubben 0f3eedd163 Do not guard a port this node is told to open to everyone (hq ADR 0103) 2026-09-22 19:44:35 +02:00
jschoubben dd6aad4a2f Give a machine port only to a port a module's container publishes, which is the only one the mesh can move (hq ADR 0038) 2026-09-22 19:44:35 +02:00
jschoubben 65187d4de0 Wait for a held node without pinning a pool connection, and give up after a bounded wait naming it (hq ADR 0100) 2026-09-22 18:31:10 +02:00
jschoubben 2e6b9d30bd Give a cascade round's hold back on every return, a body that cannot be marshalled included (hq ADR 0100) 2026-09-22 18:31:00 +02:00
jschoubben cc47330884 Reload the guard on its table rather than restart it, so a change leaves no port unguarded (hq ADR 0103) 2026-09-22 18:27:27 +02:00
jschoubben bfe4991dd7 Name every held kind of each module the flip takes in the converge preview and its digest (hq ADR 0103) 2026-09-22 18:27:19 +02:00
jschoubben 689c2c6d33 Hold a node from composing to sending, so a push composed before converge or adopt is never sent after it (hq ADR 0100) 2026-09-22 18:10:04 +02:00
jschoubben 1eff586a40 Hold the filter module to replacing the stock unit's flushing stop (hq ADR 0100) 2026-09-22 18:06:15 +02:00
jschoubben 4bb19c9e40 Give a machine port one holder: refuse ssh's, another module's and a doubled one, and release the assignment a given port replaces (hq ADR 0100) 2026-09-22 18:05:31 +02:00
jschoubben 9280513afa Refuse token issue --adopted for a converged node and point to adopt, rather than flip it quietly (hq ADR 0100) 2026-09-22 18:03:45 +02:00
jschoubben 41c300eae0 Name every kind of an untaken module's resources in the adoption envelope, not only files and containers (hq ADR 0103) 2026-09-22 18:03:03 +02:00
jschoubben c91fe1a6eb Converge only on the preview the operator saw, named by its digest, and never on an account older than 15 minutes (hq ADR 0100) 2026-09-22 18:02:30 +02:00
jschoubben ba0f44a36d Say in the converge preview that routed traffic is not previewed and is dropped unless declared (hq ADR 0100) 2026-09-22 18:01:04 +02:00
jschoubben 3dc7d0e386 Preview ssh as the derived filter admits it, and say a narrowing to the private network closes (hq ADR 0100) 2026-09-22 18:00:53 +02:00
jschoubben ade6b2bfb6 Declare nftables before the guard's table, so a node joining adopted without nft can load it (hq ADR 0103) 2026-09-22 17:59:32 +02:00
jschoubben a2dfaaf4d1 Guard only packets addressed to this machine, and order the guard's unit before the network and against shutdown (hq ADR 0103) 2026-09-22 17:59:09 +02:00
jschoubben a82bfb41f2 Leave an IPv6 loopback mapping out of what a container publishes, reading ports from the end (hq ADR 0100) 2026-09-22 17:58:50 +02:00
jschoubben 8db66e9532 Derive the guard from taken modules only: their published private-network ports and their manifests' guards (hq ADR 0103) 2026-09-22 17:58:37 +02:00
jschoubben 28b7fb81ba Write the registry's trust into the runtime's file and reload the runtime instead of restarting it; prefix reload-on like restart-on (hq ADR 0102) 2026-09-22 17:51:38 +02:00