Commit Graph
13 Commits
Author SHA1 Message Date
jschoubben a92c11be12 Resolve: one un-hostable assignment no longer refuses the whole node
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.

Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.

assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.

Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:37:41 +02:00
jschoubben 33fd28ffa6 licences: deliver the refresh token by the ordinary sealed path, not a bespoke envelope
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.

Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.

  - refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
    columns; internal/secrets/atrest.go is retired (nothing else used it).
  - the licence records its manager as (node, module); KeyFor delivers the refresh token to
    the manager holder and the access token to consumers, disambiguated by module so the two
    can co-locate. Accept and the reseal skip the manager holder.
  - the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
    module can re-seal a rotated refresh token with no private key of its own; the
    declaration tolerates its empty pre-adoption secret rather than refusing.
  - SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.

The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:08 +02:00
jschoubben 87c193b202 Resync hq ADR references 0044-0054 -> 0039-0049 after the hq record reconciliation 2026-09-05 12:47:13 +02:00
jschoubben 9b7ba2e20c identity: a consumer's identity fits the tightest backend, via a slug (ADR 0054)
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.

Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:51:39 +02:00
jschoubben 8174f5c41e plan is the send without the sending, so it allocates
Confining allocation to the push path took `plan` with it, and `plan`
belongs on the other side: it is a person asking what a push would do to
one named machine, so the port it shows and the secret it seals must be
the ones a push would use. Both are kept once chosen, so showing
numbers a later push would replace answers a question nobody asked.

Caught by the lab: composing the real modules stopped producing
postgres's sealed superuser, because nothing had minted it and the
read-only path correctly declined to.

The line is not question versus command. It is a person asking once
about one machine, against the mesh asking continuously about all of
them — the second is what hung, and the second is what reads.
2026-09-01 21:32:57 +02:00
jschoubben 4d6ec5b10c One object store, not two
object-store.json and minio.json described the same thing: same image,
same provision at the same scope, same provisioner. Not two
implementations a person could choose between — one module written
twice. Assigning both to a node would have collided on `s3-bucket`.

It exists because it was written first, to pair with photos.json for the
README's worked edge, and minio.json was the fuller version of the same
module written later. Nobody removed the first.

The pair test keeps its point and now reads the surviving one. Checked
across the rest: this was the only duplicate.
2026-09-01 21:06:17 +02:00
jschoubben 0d975a051d Asking what the mesh would send must not change it
`status` hung. It composes a declaration for every node to answer *is
this machine running what I would send it*, and composing one assigns
each module a machine port — so the question wrote to the database, and
wrote to the same rows as the machine it was asking about.

`port_assignment` is unique on (node, machine). Two transactions
inserting the same port do not race, they queue: the second waits on the
index until the first commits. A status polled every two seconds while a
node applies is two writers on those rows, and the poll stopped
returning rather than returning something wrong — which is the better
failure of the two, and still a failure.

The latent version of this was there before anything polled: two
compositions running at once could both allocate.

So allocation belongs to the send path alone. The mesh chooses a port
when it commits to sending one; every other caller reads what was
chosen. A module with nothing assigned has never been sent, which is
precisely what "waiting" means — the read needs no number to be right
about that, and inventing one would make the answer worse.

Named rather than passed as a bare bool: at three call sites, `true` and
`false` say nothing about which of these two things is meant.

Checked by the lab, which now polls status throughout an apply.
2026-09-01 20:05:08 +02:00
jschoubben c67f836185 The mesh may only move a port it actually publishes
The lab caught this: a module declaring a port and running no container
had its rule set opened on 20000 while its service sat on 9101. The
firewall reported success and blocked the thing it was told to admit,
which is the precise failure the filtering comment warns about, arrived
at from the other side.

Assignment was applied to every declared port. But a container's mapping
is the thing that translates, and where there is none the software binds
what it binds — the mesh choosing a number does not move the service, it
only makes the mesh wrong about where it is.

The declaration side already knew this: publishedOn rewrites container
ports and nothing else. Filtering did not, so the two disagreed about
the same fact. MachineSide is now the one derivation both follow.

It also fixes a second case nobody had hit yet: a mapping the manifest
wrote itself, like the mail system's 7080:80. That is passed through
untouched when composing, so assigning it a machine port would have
opened a rule on a port the container does not publish. Either side of
such a mapping now names it, and the host side is the answer — a module
may read `listens` as what its software binds or as what the machine
exposes, and both readings want the same number.

Recorded either way, assigned or not: the map means where this module's
port is on this machine, and every reader needs that answer regardless
of who chose it.

Tests bite — making it always assignable reproduces the lab failure.
2026-09-01 19:31:05 +02:00
jschoubben 1f5b70a995 The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.

The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.

An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.

Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.

A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.

Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
2026-09-01 17:52:53 +02:00
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00
jschoubben a18c3b9d13 needs is now own-secrets, named for whose it is
It sat beside `secrets` — where a *provision's* credential lands on a consumer.
Both were name-to-path, both held something secret, and the names
distinguished them not at all. Reaching for the wrong one parsed cleanly and
failed somewhere else entirely, which is the shape of fault this whole design
exists to prevent, sitting in the manifest format.

The axis that separates them is not how secret they are — both are — but
whose. `secrets` is keyed by the provision it is for and belongs to a
relationship with another machine. `own-secrets` is keyed by a name the module
chose and belongs to nobody else.

A manifest using the old name is told the new one rather than refused with
"unknown field": whoever wrote it knew what they meant, and the mesh knows what
it is called now. An invented key is still refused as one rather than guessed
at.

Found by auditing the 19 manifest fields for whether any could be mistaken for
another. This was the only pair that could — and while checking it, a second
instance of the same collision turned up one layer down: `Manifest.Needs` and
`Resolution.Needs` were different concepts sharing a name in Go. The rename
separates those too.
2026-08-31 13:47:21 +02:00
jschoubben 6eeb10f066 A command opens each store once, not once per machine
Working out what a machine should be reaches the identity context for its
certificate and the licence context for its model access. Both were opened —
and waited on — inside functions called for every node in a push. Two machines
hid it. Fifty would be fifty connect-and-wait cycles for data that does not
change while the push runs.

So a command holds what it has open, and passes it. Each context is opened on
first use rather than up front, because most commands need one and paying to
reach three would be the same waste from the other side.

The contexts stay separate, which is the point: this is one struct holding
three connections to three databases, not one connection to a shared one. No
context reaches another's store, and each still holds only its own credential
(novox/hq ADR 0008).

A pure move again — the gate is green before and after, and no test changed.
2026-08-31 13:35:39 +02:00
jschoubben 8623613704 Split main.go along the seams it already had
2,769 lines and 59 functions, holding command parsing, store opening,
resolution, the board, rotation, licences, builds and status rendering.
Nothing in it was wrong. It grew because appending was always the cheapest next
step, and no single edit was the one that should have been a new file.

That is exactly how novox/hq ADR 0001 records `hal/sdk` reaching 155 files and
34,636 lines — "containing code from every context", with each addition
avoiding a cycle and none of them the mistake. This is the same shape at 8% of
the size, which is why it is worth doing now rather than noting.

Eight files, along boundaries that already existed: what a machine is; the
private network; the catalogue; working out what one machine should be; sending
it; builds; the three questions; and reaching each context's store. main.go
keeps what a main is for — parsing arguments and dispatching.

A pure move. No behaviour changed, no test changed, and the gate is green
before and after — which is the only thing that makes a refactor this size
safe to do in one commit.
2026-08-31 13:30:54 +02:00