0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.
A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.
So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.
No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.
Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.
Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.
Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.
Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.
The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.
0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.
And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.
The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.
A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.
What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.
6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.
The filter fault is fixed. The eleven hours are the measurement, not the bug.
Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.
Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.
Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.
What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.
So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
Both found by converging the control node — the first machine with a firewall to
flip, since the two before it had none.
143: the preview and the flip both say the found firewall is disabled. The node
reported converged, 372 resources applied, nothing failed, and ufw is still
enabled and active. A converged node's declaration carries no resource that would
disable it; the sentence is printed by the command and nothing implements it.
144: ufw was never what filtered the traffic that mattered there. Fifty forwarded
openings converged through it had matched zero packets, while a chain the
predecessor installed in the container runtime's pre-accept hook did the work —
in memory only, recreated by nothing. The mesh's filter now covers that path, so
the machine no longer depends on it, but the chain remains and is the only thing
refusing the bus and the registry, which the design requires reachable from
anywhere so a machine can enrol before it has a private address.
Measured: the host is a binary somebody copied to four machines, owned by no
package and built by nothing, while the controller, catalogue, builder and vault
are container images publishing no ports at all. Same language, same project,
same kind of work, delivered two ways — and the difference is not a judgement
about either, it is that images are the only delivery that works.
What it costs: genesis must raise a container runtime before the control plane
can exist; updating the control plane goes through a registry the control plane
runs; a host change cannot be rolled out at all; and compiling the language the
mesh is written in is not a capability of the builder, so the controller is built
from a hand-written Dockerfile — the incantation the bundle toolchain exists to
abolish.
Third-party software stays a container: the store, the registry, the broker are
somebody else's build. The container runtime stays on the machine for modules.
What changes is that the control plane no longer needs it to exist.
The receiving half is already built and tested (ADR 0141). Staged: compile Go, an
artifact names its target, deliver a binary, the host first, then the rest, genesis
last.
The record claimed a version reaches a machine as an ordinary archive with
nothing new needed. Two things it needs do not exist: no toolchain can compile
the host (the list is typescript and python, and the control plane, also Go, is
built as an image from a Dockerfile instead), and nothing interpolates a built
version into a resource path, so nothing can ask for .../versions/<version>/.
The decision, the options weighed and every consequence stand — the host half is
merged and tested. What was understated was the cost, so it is corrected in place
and dated rather than superseded.
The supervision was already right — a clean exit means the host stood aside and
the launcher runs what is on disk, failures are counted, and a rollback happens
at the limit. Two things made it dead code: nothing told the running host a
successor was waiting, and the rollback resolved its known-good version through
pacman, which no machine here uses and which two of three operating systems do
not have.
Keeping a version rather than a path was the clue. Versions live side by side in
directories named for them; the newest runs; the running one stands aside between
reconciles; a reconcile that completes records itself and retires what is older
than its predecessor; rollback starts that predecessor. No new resource kind and
nothing new on the bus — an archive already fetches by digest, and the path
written is never the path executing.
Answers issue 142.
A host change merged yesterday reached no machine without a person copying a
file. The host is not a build target, no declaration delivers it, and the half
that recovers from a bad host — noticing the executable changed, a known-good
record, a launcher that rolls back — is written, tested and called by nothing.
All four machines run a byte-identical hand-copied binary that no package owns
and no record names, so nothing can say a machine is behind.
Found because ADR 0140 needs the machine to report a new fact, and merging that
could not roll it out.
Reading a converged machine's rendered rules showed the cause: the chain blocks
everything passing through and then allows the machine's own containers back by
listing their address ranges. 0137 made that list typeable and 0139 tried to
generate it; both refined a list that should not exist, because the mesh has no
position on a container reaching outward. Constrain what arrives from outside,
allow what did not, and let the machine report which links face outside — one
fact instead of a list. Ports keep following the modules unchanged.
The records check now allows one record to supersede several, and stops
requiring a withdrawn record's own citations to be live.
Both follow from the same rule the mesh is built on — a node's configuration is
composed from the modules assigned to it. Reach was settled separately by the
filter, the proxy's names and the certificate authority, so "this must not be
public" could not be written; it becomes one value on the assignment that all
three read. And the forward chain consulted two constants plus a typed list
although modules already declare their networks; it now forwards what they
declared, with the host rendering the addresses it allocated.
Found preparing the control-node's convergence. Reach is settled independently by
the filter, the proxy's names and the certificate authority, so "this must not be
public" cannot be written and a public certificate is obtained regardless. And the
forward chain allows two hardcoded ranges plus a typed list, though the mesh
already knows which networks exist because its own modules declared them — a range
wide enough to keep four of them would have forwarded two predecessor leftovers too.
One container restarted 2286 times over five days while the mesh reported the machine as doing what it
was told. Its overlay address was five days out of date: the host compares a container by a digest of
its spec, and the mesh's names were not in it, so a container whose image and files never changed was
left alone holding a name that no longer resolved. Forty-eight others were current only because
something else had recreated them.
The same fault as issue 045, in the field that was left out. Resolved by putting the names in the
digest.
0120 was already accepted; what failed the check was that it rests on 0112, still marked proposed —
and so do four designs. The decision stands: a definition names no node, no mesh and no host path, and
everything a module needs is a requirement the mesh resolves.
Accepting it makes the gap visible rather than hiding it, so issue 134 states it. 0112 says how it is
checked — 'a catalogue test finds no domain name in any definition value' — and there is no such test.
Asked by hand: seven modules name this installation in a value the mesh acts on, and eight mention a
public name in prose nothing reads. The two are not the same fault and the fixes differ, which is why
the issue separates them rather than counting to fifteen.
Two of its statements are built — a version preparing its state, and the mesh saying what it applied —
so the document is in-progress rather than proposed, and names the code that owns them. Resting on a
live record rather than a superseded one: 0127 was replaced by 0131.
What this exposes is pre-existing: it also rests on ADR 0112, which is still proposed, and a document
that is not itself proposed may not. ADR 0120 has rested on it the same way for a while. Accepting or
superseding 0112 is a decision, not a cleanup, so it stays visible in the check rather than papered
over.
ADR 0135 made a step something the mesh derives for any module that prepares its state, which turned
ADR 0052's reach into a fault: a module whose database is briefly unreachable would stop every module
declared after it on that machine — the fault issue 011 already removed for every other shape, and
the reason the catalogue migrates itself at start rather than in a step.
A step now stops the rest of its own module and nothing else; an action still gates the machine,
because genesis is a row of them and they belong to no module. What was not attempted is reported as
skipped rather than left to be inferred from silence.
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.
And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.
0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.
0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.
Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.