0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An authority inside the mesh is reached at <machine>.internal, so its own
certificate must be issued for that name — and it is the one module that cannot
be told its name by a binding, because it provides rather than requires. Written
as a literal it would be one deployment's machine name in a manifest, which is
what ADR 0056 exists to remove.
${machine:name} and ${machine:at}, beside the bound values and refused the same
way. An address the machine does not have is named here rather than discovered
later as a certificate nobody can verify.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node
path diverging from the cross-node one.
The mesh works out what a provider on ANOTHER machine serves by walking that node
— reading its manifest with that machine's port assignments, then settling the
result with that node's settings layers — before offering it to a consumer. A
provider on the consumer's OWN machine never passes through that walk, so every
step of it had to be repeated in resolve.go's servedHere and declaration.go's
here(). 038 repeated the port. Nothing repeated the settling.
So a served value the operator supplied reached a co-located consumer as the
manifest's empty default. On the ADR 0056 anchor that value is an internal CA's
root: step-ca and route-proxy on one node, route-proxy's binding carrying
root: "", an empty CA bundle written, a silent fall back to the system trust
store, and issuance stopping with nothing saying why. The same step-ca on another
node would have worked.
The second is the mirror direction. gitea declares a bare container port 3000 and
the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said
3000 — so the proxy beside it dialled a port nothing listens on and answered 502.
038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS
a provider about itself. The redirect uses the CONTRIBUTING module's assignment,
because the port is the workload's, not the proxy's; a contribution carried here
from another machine is left exactly as it is, its port being that machine's to
assign.
Both are settled in Declaration, which is the first moment the machine's ports and
the provider's settings both exist. That also removes an order dependence: the
resolver built its same-node needs mid-walk, from whichever modules had been
chosen by the time the requirement came up and in whatever order a map iterated,
so what a co-located binding carried depended on the order somebody happened to
assign things in. Re-deriving from the finished closure does not.
servedHere keeps its job — deciding whether a same-node provider serves anything
at all, which is what makes the need exist — and now says that its values are
provisional.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An empty label composed nothing, so a module served at the bare domain (a node's
own site) had to keep a full name — the one route the label model could not
express. The zone-file convention '@' now composes to the public domain itself,
no leading dot, so the apex is a label like any other. Test added.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.
Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.
Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.
And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.
novox/hq 02-DECISIONS/0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A co-located consumer of a `from: mesh` provision was told the port the
provider module DECLARED, not the host port the mesh assigned and published
it on. The same-node served facts are settled while resolving (servedHere,
here()), before a bare `ports` mapping is assigned its host port, so they
carried the declared number; only the cross-node path re-derived them after
assignment. So the provider was published on <node>.internal:<assigned> while
its own-machine consumer dialled <node>.internal:<declared>, where nothing
listens — the ordinary small-mesh case, and the one the fix for issue 018
(announce the same-node provider at all) left one promise short of kept.
Redirect same-node needs to the machine's assignment in Declaration, where the
port map is known, exactly as plan.go already does cross-node. The publish bind
is unchanged (all interfaces, scoped to the mesh by the listen's firewall rule);
only the announced port is corrected.
novox/hq 04-ISSUES/038
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container with `network: host` was skipped when the mesh injects its
`<node>.internal` names, on the belief it "shares the machine's hosts file
already". It does not: `docker run --network host` still gives the container
its own /etc/hosts (localhost and its own id only), so every internal name the
mesh wrote is invisible inside it, and a client that dials one gets EAI_AGAIN.
This surfaced with the first host-network consumer to dial a provider by the
`.internal` address the mesh hands it as `${bound:...:at}` (the model-usage
store reaching its postgres). The remedy is the same `--add-host` every other
container already gets — the runtime accepts it with `--network host`
(verified against Docker) and mesh-host emits it for any network mode.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.
Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.
- refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
columns; internal/secrets/atrest.go is retired (nothing else used it).
- the licence records its manager as (node, module); KeyFor delivers the refresh token to
the manager holder and the access token to consumers, disambiguated by module so the two
can co-locate. Accept and the reseal skip the manager holder.
- the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
module can re-seal a rotated refresh token with no private key of its own; the
declaration tolerates its empty pre-adoption secret rather than refusing.
- SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.
The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A consumer that requires a provision whose serves names no consumer key (redis-cache, amqp)
contributes no payload, but it still ASKS for it. ContributionsFrom keyed 'asks' on
contributions alone, so such a consumer's grant got From='' — read as withdrawn — and the
provider never created its account. redis-cache consumers (e.g. baserow) were silently
unprovisioned, tolerated only by their embedded fallback. A module asks iff it still requires
the provision, whether or not it hands anything up. Regression test added.
Found by the lavinmq AMQP provider bed (given:[] for a require-only amqp consumer); fix
lab-proven green there.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The provider-seal-key gate: on a node with two modules requiring the same provision (baserow
and letta both consuming postgres), sealedFor matched a need by provision NAME alone, so a
file's ${secret:X} placeholder took whichever consumer's sealed credential came last in
r.Needs -- the OTHER module's password. baserow was handed letta's password and could not
authenticate. The secrets:-map delivery path already guards this (For == m.Module, novox/hq
04-ISSUES/022); the ${secret:...} placeholder path did not. Added the same guard.
Also dedups the contributions file: when provider and consumer are co-located, grantsFor
enumerates the same-node consumer, so a consumer was emitted twice into the provider's
receives file (once full with its grant, once partial). The m.Contributes loop now skips a
(provision, module) the grants loop already carried; non-grant contributions (routes) still emit.
Regression test added: two consumers of one provision each get their own credential. Proven
end-to-end on a two-node lab install (mesh-lab assigned-two-node-db): baserow and letta on one
node, substrate on another, each authenticates with its own minted password.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
04-ISSUES/036: the media stack is several modules that must share the
library and download directories on one machine, but the manifest could
only say "a directory I own". Six modules each declared the same paths as
their own resources, and the resolver's duplicate-owner refusal — right
in general — would refuse the stack's only sensible assignment the first
time two of them landed on one node.
Add an `accesses` field: a pre-existing, operator-owned path a module is
granted use of but does not own (novox/hq ADR 0051). Distinct from a
`directory` resource on every axis the host acts on — the mesh creates,
chowns and reconciles a directory; it mounts an access and owns nothing.
An access is not a resource, so it never enters the duplicate-owner map
and several modules may name one path with no conflict. What is refused
is the contradiction: a path one module owns and another accesses.
Rendered into the declaration as an `access` resource, before the
container that mounts it, so the host can find it present or refuse
clearly. Unit tests cover co-resolution (the exact 036 case), the
unchanged owner-vs-owner refusal, the owner-vs-accessor refusal, and
access validation.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.
Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
listens.from was a manifest constant — one value for every node a module runs
on. Now a per-node setting overrides it: {"expose": {"5432": "anywhere"}} makes
postgres public on the machine it is set for while it stays from:mesh elsewhere,
and the firewall (ADR 0050) is computed from the effective source. Exposure()
validates it — a port the module does not listen on, or a source that is not
mesh/anywhere/machine, is refused rather than reaching nothing; UnusedSettings
knows 'expose' is a real destination. Tested: default mesh, setting opens it to
anywhere, bad settings refused.
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.
The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.
An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.
Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.
A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.
Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
novox/hq 04-ISSUES/025. Every image reference in every example module
was sixty-four zeros — eighteen of them across five modules. Each
parsed, resolved, and composed into a declaration a host accepts, and
none could ever have started: the machine reaches `docker pull` and
stops. That is why those modules were written and not running, and no
check saw it because every check passed.
The host validates the shape of a reference and nothing more, which is
correct: verifying a digest exists means reaching a registry, and that
is the one thing a host must never have to do. So the last place that
could catch this is the wrong place to try.
The guard therefore sits where a declaration is composed, not where a
manifest is parsed. A file in a repository is allowed to await a pin —
the design already says the manifest in a repository names artifacts
while the manifest the mesh holds names digests, and the bundle works
exactly that way. What must never happen is a placeholder reaching a
machine, and composing is the last moment before one does.
Twelve third-party images resolved to real digests without pulling
anything, which is also the mechanism the open issue needs. Two
discoveries came free: mailu publishes to ghcr rather than Docker Hub,
so seven references named repositories that do not exist at all; and it
renamed roundcube to webmail, so that one would have failed even with
the right registry.
What stays a placeholder is the mesh's own provisioner images, which
genuinely have no digest until built and pushed — the bundle's problem,
legitimately unresolved here. The stand-in consumer now stands in with
a real image rather than an invented one.
A contribution that is not a credential grant — a module offering
something to another on its own machine — has nobody to be identified
to, and was carrying an empty `as`. A field that is always present and
usually empty teaches a reader to ignore it, including when it is not.
novox/hq 04-ISSUES/023. A consumer was given its password, the address,
the port and where its credential lives, and still could not connect —
the user name was invented by the provisioner and recorded nowhere, and
the rest sat in a JSON binding that a program reading KEY=value cannot
use.
Both halves have the same cause: the mesh knew something and did not say
it.
**Who a consumer is, said once.** The provisioner used to derive
mesh_<node>_<module> and that string existed nowhere else — not in the
control plane, not in the binding, and above all not at the consumer,
which has to present it. Now the mesh derives it once and sends it to
both ends, so they agree by construction rather than by two conventions
that were the same on the day they were written. The provisioners refuse
to invent one if the mesh says nothing, because falling back to a name
of their own would create a role the consumer would never guess and
everything would report success.
**Bound values reach the file that needs them.** ${bound:provision:key}
is the symmetric twin of the sealed placeholder, and simpler: these
values are not secret, so the control plane fills them in before sending
and the host gains no field and learns no format. It stays
name-agnostic — at, as and from are true of any provision, and every
other key comes from what the provider said it serves.
The asymmetry it removes was backwards. The secret is the hard case,
because the mesh must not be able to read it, and the secret was the
part that already arrived.
Keycloak and Gitea now produce complete connections, asserted from the
manifests on disk rather than from fixtures: every part filled, no
placeholder surviving as a value, and the password still a hole only the
host can close. Three faults injected, each caught.
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.
The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.
Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.
Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.
Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.
Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.
Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.
Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
The gap that stopped keycloak and gitea from starting. A granted
credential arrives as a file whose entire content is the password, which
is what a program reading a password file wants — and most programs do
not read one. They read KEY=value, or a JSON document with the token at
an attribute inside it. A module in that position could be handed the
bare value or nothing, and both are useless.
The host has been able to do this all along: content with ${secret:name}
in it, sealed values beside it, substitution on the machine, which is
the only place both halves exist. Nothing filled the values in, so the
hole could be written and never closed and the host refused the file.
That refusal was correct and the feature was unreachable.
A module reaches its own secrets and the credentials it was granted —
both things it wrote in its own manifest — and nothing else. Naming
another module's is refused: two modules on one machine are as separate
as two on different machines, and letting one read the other's
credential by guessing a name would end that to save writing a file.
Filling runs after settings, which is the whole reason it sits where it
does. A setting is how a placeholder gets into a JSON document in the
first place — the desktop client that reads its token from an attribute,
not an environment variable. Before the merge that file's content is
"{}" and asks for nothing.
Tested through Declaration rather than through the helper. Three times
in this repository a test asserted on a helper while the code calling it
was wrong, and each time the injected fault stayed silent. Three faults
injected here — the call removed, the call moved before settings, and
the module boundary widened — each caught by the test meant for it.
It sat beside `secrets` — where a *provision's* credential lands on a consumer.
Both were name-to-path, both held something secret, and the names
distinguished them not at all. Reaching for the wrong one parsed cleanly and
failed somewhere else entirely, which is the shape of fault this whole design
exists to prevent, sitting in the manifest format.
The axis that separates them is not how secret they are — both are — but
whose. `secrets` is keyed by the provision it is for and belongs to a
relationship with another machine. `own-secrets` is keyed by a name the module
chose and belongs to nobody else.
A manifest using the old name is told the new one rather than refused with
"unknown field": whoever wrote it knew what they meant, and the mesh knows what
it is called now. An invented key is still refused as one rather than guessed
at.
Found by auditing the 19 manifest fields for whether any could be mistaken for
another. This was the only pair that could — and while checking it, a second
instance of the same collision turned up one layer down: `Manifest.Needs` and
`Resolution.Needs` were different concepts sharing a name in Go. The rename
separates those too.
Three manifests and the rule that keeps them apart. Serving and asking are
genuinely different roles, and systemd-resolved can only do the second — it
cannot answer a wildcard, it routes the mesh's suffix to something that can. A
module that treated them as one role could not work, which is the mistake worth
naming rather than discovering.
So `the-dns-port` and `the-resolver-configuration` are two claims. A machine
gets one of each, and two of either is refused by the mesh rather than fought
over on the machine — which is what ADR 0009's table meant by listing resolvers
beside the seat and pid 1. That table names the resource `/etc/resolv.conf`,
which is what it is; a claim is a name in the catalogue's own form, and the
catalogue refuses the path as one.
Neither module knows anything about the machine it is on, which is what lets
them be static manifests: they name `mesh0` and `127.0.0.54`, both chosen by
the mesh, rather than an address only that machine has. Not 127.0.0.1 and not
127.0.0.53 — taking either would be a module claiming something it did not say
it claims.
A service can now reflect a file another module put on the machine, written
`<module>.<id>`. The resolver has to restart when the mesh rewrites the names;
without it, it would serve the names it started with for ever, with every
machine that joined afterwards unreachable and every check passing.
Internal names are written to the machine's hosts file, which serves the
machine and not what the machine runs: a container gets its own hosts file
holding only its own hostname. So every name the mesh wrote was invisible to
the majority of things that need one — and on the machine it always worked,
which is exactly what made it easy to miss.
It was hit for real in the lab, and worked around by resolving the address on
the machine and passing it in. That workaround is now removed, and its absence
is the assertion.
A file rather than a resolver, which is the decision the mesh already made
about names and this extends rather than overturns: it works on every runtime,
needs no package and has no failure mode of its own. The stated trigger for a
resolver — names that are not one-per-node, service names, wildcards — is
still not met.
Given by the mesh, not chosen by a module: a module that listed the machines
would go stale the day one joins, and one that did not would be a module whose
containers cannot reach anything by name. A container that named its own keeps
them and gets the mesh's beside them.
Only containers, and not the ones on the machine's own network: a runtime
refuses to write a hosts file for those, and a file or a service given the
field is a declaration the host refuses outright — so getting it wrong breaks
the whole machine for something that was never about names.
The machine that most needed a firewall was the one that could not have one. A
hub is dialled by every node at other sites and needs its port open; a machine
that is not a hub dials out and needs nothing open. They are the same module,
and `listens` in a manifest is one answer for every machine that runs it — so
the machine a static answer gets wrong is the one facing the public internet.
A generator can now say what it opens, in a second interface rather than a
method on every generator: most have nothing to say here, and requiring an
empty method of each would be a cost paid everywhere for one caller.
The port is the one in the endpoint, which is where the interface takes its
ListenPort from. One source, so a rule set cannot open a port the interface is
not on. Open to everywhere and deliberately: a node at another site is not on
the private network until this port lets it on, so restricting it to the mesh
would be a rule that can never be satisfied by the thing it exists for.
And a generator that cannot say is refused rather than read as silence. Closing
a port on the evidence of a failure to look is how a machine is severed by a
fault somewhere else — and the machine it would sever is the hub, whose only
route to being fixed is the network it just closed.
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first
thing here that no machine can answer: a hosted model is on nobody's node and
is reached over the public internet, so the rule that refuses two ends sharing
no private network must not apply to it.
A licence is a named thing and the name is the operator's — *the personal
account*, *the organisation's* — because the whole point is saying which one a
given consumer uses, and an anonymous credential hanging off a provider cannot
be said. Many to many, so deliberately not a claim: two machines sharing an
account is ordinary rather than a collision.
Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to
each holder, discard the plaintext. With the consequence stated rather than
hidden — a holder recorded after the key was supplied has no key and the mesh
cannot make one, so it is refused by name with the remedy, not silently handed
an empty file.
Refusal is felt, as the record warns: a mesh holding three ways to reach a
model refuses every consumer that has not chosen. So the refusal names the
candidates and the exact command. Being right is not the same as being usable.
Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction
rather than a declaration — remain gaps. Half-building them would put a
conditional in the declaration language, which is what ADR 0024 says plainly to
avoid.
Its own context, with its own store and its own credential: a licence is a
different aggregate from anything inventory owns, and it refers to nodes by
name because that is what crossing a context boundary may carry.
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.
One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.
The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
A binding was skipped when the provider turned out to be on the same node,
reasoning that a file saying "it is on this node" is a fact nobody needs. That
is right about the location and wrong about everything beside it: a binding
also carries what the provider said a consumer must know, which is the port,
and a consumer cannot invent that.
A build machine sharing a node with the registry it pushes to sat in a loop
saying it could not read its own binding. Nothing was wrong with the machine,
the module, the credential or the provision — the file was never written, and
the absence looked exactly like a mistake in the module.
The original intent is kept where it was right: a provision whose provider said
nothing a consumer must know is still not written. A shell is answered here and
there is nothing to say about it. A registry is answered here and the port is
still unguessable.
The address is this machine's name on the private network, or loopback when it
has none — a machine off the network still reaches itself, and a name nothing
resolves is worse than an address that always works.
The host does not sort — order is stated (novox/hq ADR 0005) — so the order the
mesh writes down is the order a machine applies. Certificates, credentials,
bound files and the rule set were appended after a module's own resources, so a
service or container that depends on one was applied before it existed.
It failed and the next reconcile fixed it, which is why nothing caught it. A
fault that repairs itself on the second attempt is worse than one that does
not: what gets remembered is that it works.
Nothing the mesh computes depends on a module's resources, so putting all of it
first is unconditionally right. Merged after the computed-resources branch,
which replaces a module's resources wholesale and would otherwise discard them.
A rule nobody derives is a rule somebody keeps in step by hand, and five HAL
manifests carry a `scope:` key that reads as a restriction and restricts
nothing. Both halves are closed here.
Manifests are parsed strictly. An unknown key is refused, which is the
discipline the host's declaration parser has always had; `scope:` survived
because nothing rejected it.
A module says what it listens on and who may reach it, and saying from where is
required — a rule with no source is open, and must say so rather than appear to
restrict something. The mesh gathers every assigned module's ports, widens
where two overlap, names every module that wanted each one, and renders one
nftables file per node. What no module declared is closed.
Three things it deliberately does not do: it writes no forward policy, because
what a machine routes is the container runtime's business and dropping there
stops every container on the node; it never flushes the whole ruleset, only
its own table; and it carries no command to load itself, because the link may
not carry an action. A service declares `restart-on` the file instead, which is
the shape that rule leaves.
Also fixes a fault the lab found: certificateFor asked where every node is
without the catalogue, so nothing resolved, every machine looked like it was on
no private network, and every certificate the mesh was asked for was refused
with a reason that was not true. Asking that question without the catalogue is
now refused rather than answered wrongly.
08-connectivity keeps two authorities apart on purpose: a public one for
names the outside world reaches, and the mesh's own for names only the
mesh knows. Nothing implemented the second, so anything between machines
was plaintext or trust-on-first-use — which the design refuses everywhere
else.
A node now generates a fourth key at enrolment and reports the public
half. A fourth, because a key used for two purposes is one rotation away
from breaking the other: the identity key signs messages to the mesh and
would do for TLS, and reusing it would mean rotating a node's identity
every time its certificate is replaced.
**Nothing secret travels and nothing is sealed.** A certificate authority
says "this name belongs to the holder of this key", so the mesh signs a
public half it cannot use, and the certificate it issues is public. A
module asks for one and is given the certificate and, if it wants,
the mesh's own — the private key is a path to a file the machine already
has, the same arrangement the private network's key uses.
Asserted by verifying rather than inspecting, because a certificate that
parses and does not chain fails at the moment something connects:
- what the mesh issues verifies against the mesh, for the name asked for
- the name is in the subject alternative names, since a certificate
carrying it only in the common name is refused by every modern client
- it certifies the key the node generated and no other
- another mesh's certificate does not verify, which is the whole point of
two authorities being separate
- the authority cannot sign another authority — one that could is one
that can be delegated without anybody deciding to
- two control planes starting together agree on one authority, or a mesh
has certificates half its machines refuse
Certificates last ten years, which is a choice: a short life needs
something to renew it, and a renewal that fails silently is a mesh that
stops trusting itself on a date nobody wrote down. What makes one
replaceable is that the mesh reissues on demand, not that it expires.
Found by testing removal, which is the half nobody tests.
A grant was emitted for every secret the mesh held, whether or not the
machine still asked for it. So a consumer that was unassigned kept
appearing in its provider's manifest — and the provisioner's rule about
removing what nobody asks for can only fire if the mesh stops asking. The
login would have stayed live for ever, and nothing would have said so.
Skipped where the declaration is built rather than where grants are
gathered, so the rule holds whoever gathers them. No credential file is
written for a withdrawn consumer either, or the provisioner would find a
file its manifest does not mention and have to guess what that means.
The secret itself is deliberately kept. It is sealed and unusable to the
mesh, and a machine that comes back gets what it had — what withdraws the
login is the manifest, which is the thing that reconciles.
Two things, both found by trying to write a real postgres module and
discovering it could not be said.
A database has a superuser password, a broker an administrator, a
registry an account. None of them is *for* anybody — they are not the
credential a consumer is given, and the mechanism that hands those out
has a consumer in the middle of it. So a module may declare what it needs
and where to put it, and the mesh generates one per node, seals it, and
reads it no more than it reads any other.
Per node, deliberately: a module running on three machines has three
passwords. One in the manifest instead would put the same secret on every
machine that ever runs it, in a file anybody can read, for ever. Made
once and kept, or a running database would be handed a password it was
not started with; remade when the machine's sealing key changes, like
everything else sealed here.
A need declared and not made is refused rather than skipped, because a
module whose own credential is silently absent starts, fails to
authenticate, and the reason is three layers from the machine reporting
it.
And the provisioner can watch. That is what lets it be a module rather
than a binary somebody places: run once, it needs invoking after every
declaration by a timer or a unit wired to a file; watching, it is an
ordinary long-running service the host already supervises. It polls
rather than watching the filesystem, because the host writes atomically —
the file is replaced, so a watch on the path stops seeing anything after
the first replacement, and a watcher that silently stops working is worse
than a poll. Credentials are compared by digest and never held: this runs
for as long as the machine is up.
resolve.go had grown to 796 lines doing four jobs: working out what a
machine should run, applying settings, collecting contributions, and
placing credentials. They answer different questions — the first is
"what", the rest are "what does that look like as resources" — and one
file doing both is how a thing starts becoming the kernel everything
imports.
Prompted by looking at why HAL's shared library became unmaintainable.
Measured while here, and the shape is the inverse of that one: the large
packages import nothing internal, and only inventory and link compose. A
change to module resolution cannot reach connectivity, because
connectivity does not import it.