Two sibling branches resolve a provision answered by the consumer's own machine.
The one for a node-scoped provider falls back to loopback when the machine is on
no private network, with a comment saying why and a test holding it. The one for
a mesh-scoped provider passed node.At straight through, and nothing noticed
because nothing had yet composed a host out of it.
The mesh's own artifact store is mesh-scoped and sits on the same machine as the
builder that pushes to it. Give the builder the address from its binding and it
gets MESH_REGISTRY=:5000 — a name with no host, written into its environment
without complaint. It surfaces much later as
":5000/mesh-tools/build" is not a valid repository/tag
which is a message about a tag for a fault in how a binding was resolved, on a
machine several steps from the decision.
A machine off the private network still reaches itself, which is what the
neighbouring branch already said. The test fails without the fix, showing the
empty address rather than only the symptom.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The check read "declares no resources" as "runs nowhere", and those are not
the same. The private network declares no resources either — the control plane
computes them when it composes a machine's declaration — and it is assigned to
every machine that has to reach another one. Refusing it stopped a four-machine
bed at its first assignment.
The signal is narrower: it builds an artifact and places nothing. Made a
function of its own, because a judgement with a wrong answer this expensive
should be testable without a database — nothing guarded it, which is how it
shipped.
The floor allowed ssh from the mesh's addresses, and from everywhere on a
machine that faces outward. On a machine the mesh knows no addresses for it
emitted neither — so the chain dropped by default and ssh was simply shut.
That is the first machine anybody adopts: reached over the network, with the
port needed to fix it closed by the act of adopting it. Found by reading the
rules off a live machine rather than trusting the generator.
Two faults, opposite directions, both in issue 047.
There was no forward chain, on the reasoning that dropping there stops every
container the runtime allowed. The first half is true; the conclusion was not. A
published port is redirected and then forwarded, so it never reaches the input
chain — the firewall was silent about the ports most worth protecting. The way
through is the one the system being replaced already used: deny by default, then
allow the runtime's own networks explicitly. A forwarded rule matches what the
client originally asked for, because the destination has been rewritten by the
time the chain sees it.
And ssh is now a floor nothing derives. Every other line comes from what is
assigned, which is the point — but a mesh part-way through adopting a machine
has been assigned almost nothing, so what it computed was a chain that shut the
port used to fix it. From the mesh always; from outside on a machine that faces
outward, because that is the way in when the private network is what broke.
Rehearsed on three machines: a docker-published port declared mesh-only is now
reachable from inside the mesh and refused from outside. Before, it was
reachable from both.
The image every module in the scripted toolchain is compiled on top of is
registered as a module so the mesh can build, version and depend on it. It is
not one: nothing about it belongs on a machine. Assigning it succeeded, the
machine was sent a declaration containing nothing of it, and everything
reported success — the operator had said run this here and the mesh had agreed
to something it cannot do.
A module whose resources are worked out per node is asked about separately, so
it stays assignable, which is the point of it.
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
A container naming an artifact is a module saying the mesh builds this. Until a
build publishes one there is nothing to run — and what reached the machine was
an unresolved field, which its language has no room for, so it refused the whole
declaration and reported that a container does not use "artifact". That reads
as a broken manifest. It is not broken, it is unbuilt, and only the mesh can
tell those apart.
Found by the four-machine bed, which assigns modules the mesh has not built.
A module is a repository and a path within it, and the manifest sits at that
path. This one sat in the catalogue instead, so the thing that says what the
control plane is and the thing it is made of lived in different repositories
and could drift apart with nothing to notice.
It also names an artifact it builds rather than a placeholder digest somebody
fills in, which is what lets the mesh build its own control plane.
This is how a mesh is raised: the installer carries this program and runs it
once, before anything exists, to produce the control plane from the same
repository and path every later rebuild will use. What raises the mesh is then
the same thing that maintains it, rather than a second mechanism exercised once
per new mesh — which is how often enough to rot.
With nowhere to publish, an image stays in the machine's own runtime and is
named by the digest of its own configuration: the same identity the installer
has always used for the image it carried.
The builder says what it built and the catalogue decides whether that was an
upgrade. Only the control plane knows which machines run the thing, so it is
the one that acts — and what it does is a choice somebody recorded, not a
behaviour compiled in: record that they are behind, or send it, one machine at
a time or together.
Recording is the absence of an action rather than a second path: a machine not
running what the mesh would send it is already something the mesh reports.
Defaulted to recording. A mesh that rolls out everything it builds the moment
it builds it is reasonable to want and a bad thing to arrive by default — the
first module to inherit it would be the control plane, upgrading itself out
from under the push applying it.
An event whose origin reads a container id names something no other module
can look up. The mesh already knows the answer, and a module's environment
file is a file resource, so ${machine:name} reaches it with no composer change.
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.
What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.
Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.
Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).
The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.
A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A store connection string carries a password, and the control plane took it
from MESH_STORE_<CONTEXT> — an environment variable, which is readable in
`docker inspect`, in the process's own /proc entry, and in whatever composed
it. Every other module in the catalogue is given secret material as a file the
mesh sealed to the machine and the host wrote.
That difference is what stopped the control plane from being an ordinary module
(novox/hq ADR 0067). A manifest can put a sealed value into a file's `content`
with ${secret:…}; it has no substitution into a container's `env` at all. So a
control-plane manifest could be written with the password in it, or without the
setting — neither honest. The fix is not to change the manifest format but to
let the control plane read what everything else reads: a file.
MESH_STORE_<CONTEXT>_FILE names one. Exactly one of the two may be set; both is
refused rather than settled by precedence, because whichever won, the other
would still read as the setting in force and the process would be writing to a
store nobody expects. Trailing whitespace is trimmed — a file written by a
person or by a filled-in placeholder ends in a newline, and a newline inside a
URL is rejected several layers from anything that could explain it. Leading
whitespace is left, being a mangled value rather than a habit.
The no-leak property is kept and extended: a file that cannot be read names its
path, never its contents, and the parse failure now names whichever source was
used because a variable name and a path are not the secret.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The sibling of node public-domain, and the worse one: a placement is three facts
declared together, so an invocation that said none of them took all three away —
the endpoint every other machine dials, the site, and the hub. A mesh whose hub
was placed that way has no paths left, at the moment somebody was trying to look
at it.
--nothing keeps the real case (a machine that roams and opens every path itself)
sayable, by name.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`status --json` emitted no JSON at all when a single node was unresolvable. A blocked
node is not on the private network, and a mesh whose hub is that node has no hub — which
came back through the reading as a refusal, so `status` printed nothing and `status
--json` put multi-line prose on stderr and not one byte on stdout. A machine-readable
interface that stops being machine-readable exactly when something is wrong is one nobody
can build an alarm on.
Why a machine cannot be worked out is read as data now, per machine, through the same
whoResolves the private network is built from — so this and the network agree about who
could not be resolved rather than deciding it twice. The private network failing to
compute is kept as a note beside it instead of ending the read: it is almost always a
consequence of those same refusals, and every question that does not depend on it is
still answered.
It reaches all three ways of saying it, from the one reading: the text form leads with it
because a machine here is in none of the answers below, the JSON carries `unresolved`
(always a list, never null) and `network`, and the page has a section of its own.
That also closes a silent success. A machine that resolves to nothing has nothing
computed for it, so there is nothing to compare it against and nothing it can be behind —
it appeared in no answer at all, and `status` reported a mesh where nothing could be sent
anywhere as "all doing what they were told".
statusAsJSON takes the whole reading now rather than a growing argument list, which is
what let an answer be added to the text form and forgotten here. The two are one
function's output in two shapes and must not be able to differ about what was asked.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An authority inside the mesh is reached at <machine>.internal, so its own
certificate must be issued for that name — and it is the one module that cannot
be told its name by a binding, because it provides rather than requires. Written
as a literal it would be one deployment's machine name in a manifest, which is
what ADR 0056 exists to remove.
${machine:name} and ${machine:at}, beside the bound values and refused the same
way. An address the machine does not have is named here rather than discovered
later as a certificate nobody can verify.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`node public-domain <name>` cleared the domain. It reads like a question — it is exactly
what anybody types to find out what the answer is — and it silently took every routed
name the node had. There is no output that makes up for that: by the time it prints, the
fact is gone, and the mesh cannot tell a person what a domain used to be.
The bare form reports now. Clearing is still a real thing to want — a machine that stops
facing the outside composes no names, and lab-versus-production is this one setting
(novox/hq ADR 0056) — so it keeps a way to be said, by name: `--clear`. A domain and
`--clear` together are refused rather than one of them silently winning.
The other `node` subcommands were checked. `add`, `list` and `show` write nothing they
were not asked to, so there is nothing to make consistent with.
`overlay place <node>` with no flags has the same shape — it clears the endpoint, the
site and the hub flag — and is deliberately left alone here. It is a verb rather than a
question and every caller passes flags, so the fix is a different judgement and belongs
in its own change.
The usage text gains the three forms, and `module forget`'s new flag, neither of which
it named before.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An assignment was verified by resolving the node it was made on. The verify that matters
is resolution over all of them: a module offering a mesh-scoped provision stops offering
it the moment its own node stops resolving, so an assignment could be reported as fine
while it took that provision away from every consumer elsewhere. Those consumers were
then told "nothing in this mesh provides it", naming as the remedy a module that was
already assigned — a wrong answer about a machine nobody had touched.
That is novox/hq 04-ISSUES/017's shape exactly: an action succeeds into a state its own
verify rejects, and it does so because the action's own test is not the test the verify
uses. 017's remedy was to make them the same test, and this makes them the same test.
The assignment is still kept, and that is the other half of the decision. Assignment is
not an ordering: a consumer assigned before its provider does not resolve for as long as
it takes to assign the provider, and refusing the first half of a pair would make the
order somebody types two commands in part of the mesh's rules. So `assign` and `unassign`
now name every OTHER machine that cannot be worked out as things stand, in the mesh's own
words, beside whatever they already said about this one. It reports the state and never
claims causation — saying "this assignment broke laptop" would mean resolving the whole
mesh twice and would still be a guess about which of several changes did it.
It costs a resolution per machine. Assignment is a person typing a command, and being
told which machines this just blocked is worth more than the milliseconds.
This is what a four-node raise read as "bumping a module's version broke provider
recognition". It was neither the version nor the provider: nothing in this codebase reads
the version column, every lookup is keyed on the module name alone, and a test in the
previous commit now says so. It was one machine's set of assignments, and nothing said so.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`module forget` cascaded. The settings, the module's own secrets and the ports the mesh
chose all name the module by a foreign key that cascades, so removing the row took all
three and reported "forgotten" — an action succeeding into a state its own verify would
reject (novox/hq 04-ISSUES/017). A sealed secret is not recoverable afterwards, because
the mesh discarded the plaintext when it made it.
It now reads what it would destroy, names each thing one at a time, and refuses.
`--and-what-it-holds` is how somebody says they mean it, and the removal then reports
what went — this being the only record that any of it ever existed.
Reported as "operator settings do not persist, because re-registering a module
cascade-deletes them". Half of that is wrong, and the test now says so out loud: the
upsert is on the name, so `module add` at a new version leaves the settings, the secrets
and the ports exactly where they were. The command that destroyed them was `forget`, and
a wrong belief about which command destroys data is expensive in both directions — it
sends people looking for a fault that is not there, and leaves the real one unexamined.
Checked by internal/inventory/forget_test.go, which writes all three, re-registers the
module at a new version, reads them back, and only then tries to forget it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
autocert keeps its account key at one fixed name, `acme_account+key`, in whatever
directory it is given, and reuses it for ever. That is right while the authority
stays the same and silently wrong the moment it does not. Re-initialising the
internal CA makes a new authority with a new root: it has never heard of the
account in the cache, rejects every use of it, and autocert has no path back from
that. Nothing is re-registered, no order ever reaches the CA, and issuance stops
with nothing saying why — until somebody guesses that deleting the cache directory
by hand is the answer.
Name the directory after the authority instead of sharing one between all of them:
a digest of the ACME directory URL and the root this proxy was told to verify it
with. A re-initialised CA has a new root, the mesh delivers it as a changed bundle,
the proxy restarts on that file and lands in a directory with no account in it, so
autocert registers afresh and orders again. The healing is that "is this account
still valid" never has to be asked — an account is only ever found where it is
still valid, which needs no error codes, no probe at startup and no network call
that can itself fail.
It closes a latent one of the same shape: pointing ACME_DIRECTORY at production
after testing against staging reused the staging account, because the cache had no
idea the two were different.
Trailing whitespace around the delivered root is not a new authority — the mesh
writes that file, and a newline coming or going must not throw away an account.
Old directories are left on disk, unused: they hold the only copy of certificates
that may still be valid, and this program is not the thing that should decide a
certificate is finished with.
A proxy already holding certificates orders them once more on the first start after
this, because its account moves. Free against the lab's own CA and against staging;
one issuance per name against a public authority.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A whole-mesh push refused outright the moment any single node failed to resolve:
every node was composed first, and one entry in `refusals` returned "nothing was
sent" before the send loop ran. So on the ADR 0056 lab, one module on the anchor
requiring a provision nobody had assigned a provider for — `nothing provides
"acme-ca", wanted by route-proxy` — stopped every OTHER machine from being sent
anything. Nothing converged anywhere, and the machines that went unconverged were
the ones with nothing wrong with them. The failure and the punishment were on
different machines.
This is the rule a92c11b established one level down, where an un-hostable module
stopped taking down the healthy modules beside it, applied one level up: the blast
radius of a fault is the thing that has it. A node whose declaration composes is
sent; a node whose does not is named, with its reason, and the push still ends
non-zero — skipping is not succeeding, and a command that exits cleanly having
missed a machine is a command that lies. The message now says how many machines
WERE sent, because "nothing was sent" was the claim that had become untrue.
sendTo keeps its all-or-nothing rule, and its comment now says that push
deliberately does not share it: a rotation reaching the consumer and refusing on
the provider leaves one end holding a credential the other has never heard of,
which is a real coupling between two named machines. A whole-mesh push has no such
coupling and never did.
Composing is split out of pushCommand so the rule can be tested without a broker
and a database.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node
path diverging from the cross-node one.
The mesh works out what a provider on ANOTHER machine serves by walking that node
— reading its manifest with that machine's port assignments, then settling the
result with that node's settings layers — before offering it to a consumer. A
provider on the consumer's OWN machine never passes through that walk, so every
step of it had to be repeated in resolve.go's servedHere and declaration.go's
here(). 038 repeated the port. Nothing repeated the settling.
So a served value the operator supplied reached a co-located consumer as the
manifest's empty default. On the ADR 0056 anchor that value is an internal CA's
root: step-ca and route-proxy on one node, route-proxy's binding carrying
root: "", an empty CA bundle written, a silent fall back to the system trust
store, and issuance stopping with nothing saying why. The same step-ca on another
node would have worked.
The second is the mirror direction. gitea declares a bare container port 3000 and
the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said
3000 — so the proxy beside it dialled a port nothing listens on and answered 502.
038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS
a provider about itself. The redirect uses the CONTRIBUTING module's assignment,
because the port is the workload's, not the proxy's; a contribution carried here
from another machine is left exactly as it is, its port being that machine's to
assign.
Both are settled in Declaration, which is the first moment the machine's ports and
the provider's settings both exist. That also removes an order dependence: the
resolver built its same-node needs mid-walk, from whichever modules had been
chosen by the time the requirement came up and in whatever order a map iterated,
so what a co-located binding carried depended on the order somebody happened to
assign things in. Re-deriving from the finished closure does not.
servedHere keeps its job — deciding whether a same-node provider serves anything
at all, which is what makes the need exist — and now says that its values are
provisional.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An empty label composed nothing, so a module served at the bare domain (a node's
own site) had to keep a full name — the one route the label model could not
express. The zone-file convention '@' now composes to the public domain itself,
no leading dot, so the apex is a label like any other. Test added.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Piece A of ADR 0056 (selectable issuer). A provider that serves an empty root —
public-acme, whose root already ships in the OS trust store — leaves route-proxy's
CA bundle file existing but empty, because the mesh writes it unconditionally from
${bound:acme-ca:root}. Read that as "trust the system roots", the same as an unset
bundle, instead of failing with "holds no certificate this can trust". A bundle
that holds bytes but no parseable certificate is still refused.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.
Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.
Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.
And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.
novox/hq 02-DECISIONS/0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A co-located consumer of a `from: mesh` provision was told the port the
provider module DECLARED, not the host port the mesh assigned and published
it on. The same-node served facts are settled while resolving (servedHere,
here()), before a bare `ports` mapping is assigned its host port, so they
carried the declared number; only the cross-node path re-derived them after
assignment. So the provider was published on <node>.internal:<assigned> while
its own-machine consumer dialled <node>.internal:<declared>, where nothing
listens — the ordinary small-mesh case, and the one the fix for issue 018
(announce the same-node provider at all) left one promise short of kept.
Redirect same-node needs to the machine's assignment in Declaration, where the
port map is known, exactly as plan.go already does cross-node. The publish bind
is unchanged (all interfaces, scoped to the mesh by the listen's firewall rule);
only the announced port is corrected.
novox/hq 04-ISSUES/038
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.
Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.
assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.
Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A node-scope provider that answers a requirement on the same machine and
serves connection facts (a port) but mints no credential delivered
nothing to a co-located consumer. resolve.go only built the delivering
Needed when brokered[want] was set — true only for mesh-scope providers;
a node-scope keyless provider set local[want] instead and fell through,
so knownFor saw no binding and boundInto refused the consumer's
${bound:model-access:port} file.
Deliver the served facts as a need whenever the same-node answer serves a
non-empty set, with a loopback fallback for the address when the node is
off the private network — the reachability rule does not apply to two ends
on one machine. The brokered (credentialed, mesh-scope) path is untouched.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
One registry line — `"openai": StaticKey`. The generic static-key adapter
already serves any vendor (accept is the vendor-independent seal, deliver is the
value unchanged, and refresh/identity/usage are not implemented), so a second
static-key vendor is data, not code. The shape-selection test now asserts
For("openai") reports static-key.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container with `network: host` was skipped when the mesh injects its
`<node>.internal` names, on the belief it "shares the machine's hosts file
already". It does not: `docker run --network host` still gives the container
its own /etc/hosts (localhost and its own id only), so every internal name the
mesh wrote is invisible inside it, and a client that dials one gets EAI_AGAIN.
This surfaced with the first host-network consumer to dial a provider by the
`.internal` address the mesh hands it as `${bound:...:at}` (the model-usage
store reaching its postgres). The remedy is the same `--add-host` every other
container already gets — the runtime accepts it with `--network host`
(verified against Docker) and mesh-host emits it for any network mode.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The manager module's seal now runs over the audited tweetnacl-sealedbox-js (mesh-catalog
anthropic-manager) rather than a hand-transcribed NaCl. Regenerate the fixture's sealed value
from that new seal() over the same node key pair and plaintext, so the fixture is the new
library's output. crypto_box_seal is randomised, so the blob differs; the wire format does
not. TestModuleSealedBoxOpensInGo still opens it under box.OpenAnonymous and recovers the
plaintext, proving TS(tweetnacl-sealedbox-js) to Go interop.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.
Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.
- refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
columns; internal/secrets/atrest.go is retired (nothing else used it).
- the licence records its manager as (node, module); KeyFor delivers the refresh token to
the manager holder and the access token to consumers, disambiguated by module so the two
can co-locate. Accept and the reseal skip the manager holder.
- the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
module can re-seal a rotated refresh token with no private key of its own; the
declaration tolerates its empty pre-adoption secret rather than refusing.
- SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.
The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF