0a05fbb434d3f06f3729d48e8851865fef144d4d
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
751948f0f9 |
The lab had a registry that production does not, so it tested a fiction
The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and rewrote every manifest reference — third-party ones included — to point at it. No production mesh has such a thing. So every bed proved that a machine could fetch an image from a registry that exists nowhere else, and the bootstrap problems that only appear when a machine has to fetch for itself went unfound. What replaces it is the two things that are true in the world: **Public images come from the public internet.** mesh-lab already created a NAT'd uplink for exactly this and attached it to any machine declaring `egress`; no scenario ever declared it. They do now, and third-party references are left exactly as the catalogue writes them. **The mesh's own images have no registry and never will.** mesh-control, mesh-builder, mesh-route-proxy and the per-module runtimes are built from source and exist in no registry. A machine gets them the way an operator's machine does — they are built here and loaded onto it — and is then named by the digest of its own image configuration, which mesh-host now accepts as "an image this machine already holds". `images:` therefore means only *ours*, and a third-party entry is refused rather than quietly loaded: otherwise the fiction returns one convenient line at a time. It is per-machine as well, because "everything, everywhere" was never a description of anything real — handing whole-mesh-full's union to its two 30GiB workstations would fill the disk with runtimes nothing on them will start. **The uplink and the declared gateway would have fought, silently.** A gateway container and the transit router reach the scenario and nothing else; a default route through either is a black hole for anything outside, and it beats the uplink's DHCP route on metric. So a machine with egress states the scenario's ranges explicitly — through the same gateway or transit it would have defaulted to, so the overlay-across-NAT path is unchanged — and leaves the default to the uplink. A range with no path inside the scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an ordinary private range in fact, and letting it escape would put scenario traffic on whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested, because a decision only a full raise could check is one nobody checks. The registry-reachability check the raise gained earlier is kept, pointed at the real thing: every machine with egress must resolve a name and reach the internet before the raise says it finished. Same failure it was written for — a raise that returns, an apply that dies on its first pull, an instance left a bare shell — now guarding the path that actually carries. The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It was never only for the lab's registry: the mesh has one of its own, the `registry` module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
eba436b6b7 |
Reach one scenario from the workstation, by name
A scenario is a closed address space: two raised from the same declaration hold the same addresses and never meet, which is what lets two run at once and why the lab talks to machines through the hypervisor rather than over IP. Reaching in from outside breaks that, so it is opt-in, one scenario at a time, and reversible. `connect` takes an address on the scenario's public link and writes a resolver rule answering everything under each machine's name. `disconnect` gives both back. `connected` says what is true right now, for somebody who cannot remember. It refuses rather than guessing when more than one scenario is standing — the failure being avoided is not an error but one scenario's traffic arriving in another. It also refuses when a machine's name is already answered here for something real, because connecting would point that name at the lab, and the damage would land on the real thing. Names answer with the segment address rather than the overlay one. Inside the mesh a name gives a machine's private address; from here that would need this workstation on the overlay, which is a much larger door. The segment address reaches the same machine and the same ports, which is what opening a board in a browser actually needs. Proven against a live two-node scenario: registry.internal:5000/v2/ answered 200 from this workstation, and so did a wildcard name under the same machine. Disconnect put the address back, stopped answering, and left the real mesh's own names alone. One thing measured rather than assumed: it restarts dnsmasq instead of reloading it. A reload is SIGHUP, which re-reads the hosts file and clears the cache but not the configuration — the rule was written, the reload reported success, and nothing resolved. The daemon's start time was nine days old afterwards. |
||
|
|
a516ee847b |
Adopt a third-party workload, and keep a mesh between runs
**The adoption.** Software nobody here wrote, taking its credentials the way such software does — from its environment — and needing two containers that reach each other by name. The first module that could not have been declared this morning: it needs the network shape and it needs a sealed value to reach a container's environment. Its password is accepted rather than generated, which is the whole shape of an adoption: a service that already exists keeps the credential it already has. Asserted properly — a wrong password is refused by the same database, so the passing case means something. **The warm scenario.** A mesh kept between runs and returned to, which turned twelve minutes of bootstrap into thirty seconds of restore. Off unless asked for: a run that is meant to mean something raises from nothing. Its guard fired for real during this work, unprompted — a mesh-host commit landed and it refused the stale base, naming both commits, rather than passing tests against yesterday's binary. That is 04-ISSUES/005's rule one level down. Three things the guard learned the hard way and now handles: a snapshot captures disk and not memory, so the host is restarted after a restore and asserted to have come back; the stocked image digests are worked out while raising and a restored instance never raises, so they are kept; and comparing only the repositories this run can see clears the ones it cannot, so both directions are compared. The one real bug behind five failed attempts was in mesh-host and it reported itself precisely: a network shape the language had and no host implemented. Everything else was scaffolding of mine. |
||
|
|
033ad7ec69 |
A run rebuilds what it tests, and leaves a receipt saying what it covered
The danger is not that the suite breaks. It is that nobody notices it stopped running (novox/hq 04-ISSUES/005). The harness this replaces had not built for two and a half months and nothing said so — and this suite needs a hypervisor, so it inherits exactly that: it runs when somebody remembers, and remembering is not a mechanism. So running, recording, and rebuilding are one act: - the host binary, control-plane image and builder are rebuilt from source first. The last two both parse manifests; building one and not the other left a binary eleven hours old refusing a field the mesh had just renamed, found by a full run. - a receipt lands in XDG state — outside git, because the question is whether *this machine* has run it, and a receipt in git would be a claim about everybody's machine made by whoever committed last. - `last-run` judges it and exits non-zero when it no longer counts. Three faults found by running the thing rather than reading it, each now held by a test confirmed to fail without it: - counted() passed every test while parsing nothing. The runner colours its summary even into a pipe; the fixtures were clean text that had been imagined rather than captured. A fixture that agrees with the mistake proves the mistake. - a receipt for `suite test/lastrun.test.ts` was indistinguishable from one for the real thing — 005's own symptom, rebuilt inside its remedy. The receipt now records what ran. - a tree with uncommitted work reported the bare commit, claiming coverage of code nobody can check out. Nothing else could tell: the hash is identical either way. Proven on real machines: 22/22, against all three repositories. |
||
|
|
be176bab2e |
Automate the lab registry: a sealed machine pulls by digest
Closes 04-ISSUES/009. A scenario declares `images:` by tag; the lab stocks a
registry on this workstation where there is a network, raises it inside the
scenario as scenery, and reports the references a declaration pins -- which are
the digests THIS registry assigned, and are not knowable until it is raised.
Verified in a sealed machine, confirmed by ping to have no route out: package,
service including boot state, a container pinned by digest, and an action
inside that container. Applied, idempotent on re-apply, and read back from the
machine rather than from the apply's own report. That is the first time the
container shape has worked in the lab at all, and it was the shape blocking the
substrate bootstrap.
Four faults found by running it, three of them mine and one worth keeping:
The read-back checked that the catalog endpoint answered, by looking for the
substring "repositories" -- which `{"repositories":[]}` also contains. So it
passed on a registry holding nothing, and the failure surfaced much later as a
container that could not be pulled. It now asks for each image's manifest BY
DIGEST, which is what a machine does.
A recursive push needs its destination to exist, or incus copies the source's
contents rather than the source. The data landed one directory too shallow and
the registry found nothing where it looks.
The registry writes its blobs as root through a bind mount, so the workstation
could not remove its own scratch directory afterwards. Whoever made the files
removes them -- the cleanup now runs in a container too. And a cleanup failure
no longer fails a raise that succeeded: the scenario is standing and usable,
and saying otherwise would be a false report.
The base image build did not verify that the runtime trusts the documentation
ranges as plain-HTTP registries. Writing the file is not the daemon honouring
it, and a base image that looks right fails much later, in a sealed scenario,
a long way from its cause. It is now read back from `docker info`.
|
||
|
|
4097ff92c1 |
Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated record holding their reasoning -- the four lab records are 0016, a test defends a decision is 0017. |
||
|
|
d6eef25590 |
The lab can give a sealed machine a container runtime
ADR 0046's open consequence: "the lab needs a way to place images, and the
machine it places them into needs a container runtime, which a sealed scenario
cannot install either."
The runtime half is done, and it is research 012's reframing applied literally
-- fetch at build time on a machine with a network, apply on a target that
needs nothing. `mesh-lab base build` launches a machine WITH a network,
installs a runtime, verifies it by asking the runtime rather than the package
manager, and publishes the result. Measured: ~30s to install, ~60s to publish,
~700MiB, paid once per lab rather than per scenario.
A scenario that places `runtime` or an image is then raised from that base
image, chosen rather than declared -- a scenario says what it needs, not which
image provides it. If the base does not exist it says so and how to build it.
Verified in a genuinely sealed machine (no route out, confirmed by ping):
package, service including the new `boot: enabled`, and action all applied,
were idempotent on a second run, and read back correctly. Those three had never
run anywhere but a workstation.
The image half is NOT done, and testing found why: a digest-pinned image cannot
be placed from an archive. `docker save alpine@sha256:...` produces an archive
with no repo tag, because a repo digest only exists for an image a registry
served -- so it loads dangling and a container declaring that digest reaches
for a registry the machine cannot see.
That collides with ADR 0046, which has the host REFUSE an unpinned image. Tag
refused by the host, digest unusable in the lab: there is currently no
declaration the lab can raise that exercises the container shape at all. Filed
as 04-ISSUES/009, whose resolution is a registry inside the scenario -- which is
what the real mesh does rather than a workaround for the lab.
Also fixed a weak check of my own, which is the same fault in miniature: the
load was tested with `includes("Loaded image")`, a prefix of both `Loaded
image:` and `Loaded image ID:`. So an unusable dangling load reported success
and the failure surfaced later as a container that would not start.
|
||
|
|
2243618f01 |
Draw a scenario, from the declaration and from the hypervisor
`mesh-lab diagram` renders a scenario as draw.io, from either source, through one layout — so a difference between what was asked for and what exists is a difference you can see. The shape says what a resource is and is fixed per kind. The badges say what is true about that particular one and come entirely from metadata: translation, forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The interesting properties of a network are exactly the ones with no visual consequence — a translated address looks identical to an untranslated one. For the live picture to be a record rather than a restatement, raise now writes down what it applied: a segment's kind, ranges and MTU on the link; a gateway's translation, forwardability and expiry on the gateway; inbound: deny on the machine. Every behavioural tag is written AFTER the thing works, never at creation — a failed raise leaves wreckage standing on purpose, and a picture of that wreckage must not badge translation the router never got. The pairing earned itself immediately: drawn side by side, every virtual machine held no addresses. A container's interface carries the device's name and a VM names its own, so joining them by name silently dropped one whole class of machine. Fixed by joining on MAC. Also brings tests under the typecheck gate, which caught integration timeouts being passed as a 4th argument and therefore ignored entirely. |
||
|
|
a27d861d3b |
Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a workstation, not asserted: two machines raised and addressed in 14.6s, snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no loss, and the workstation with no route into any of it. The declaration layer implements the model in full — three positions a machine can be in, keyed on forwardability; gateways carrying the address the world sees them as; both address families; multi-homing; MTU; inter-segment policy. It is validated hard because the failures it prevents are silent: a private range on a public segment produces no error, the mesh simply never forms. Public segments are refused unless they use RFC 5737 or RFC 3849 space, and a range wider than the reserved block is refused too. 33 tests, all offline. The runtime implements less than the model, and refuses the difference. A scenario declaring gateways, published ports, policy, inbound deny or place is rejected at raise with every gap named. Raising it would produce a mesh that silently lacks what it declared, which is the fault this lab exists to catch — 04-ISSUES/003, where a firewall key is declared in five manifests and read by no code. Three bugs found by review and by running it, all of one family: The readiness check truthiness-tested incusOk's return. `exec … true` succeeds with EMPTY output, so every machine reported unreachable while incus exec on it worked perfectly. succeeds() now exists so the mistake is not available, and network delete had the same bug — it counted zero segments removed while removing them. list() split instance from machine on the last dash, so a machine called home-server absorbed half the instance id and destroy found nothing. Resources are now found by the metadata they carry, never by name. restore reported success in 0.79s while the machine's agent was still starting, so the next command failed. Both raise and restore now wait for usable and say how long that took — reporting the earlier number is transport reported as effect, which is the fault the lab is being built to find. Two incus behaviours worth recording. Its CLI reads a YAML definition from stdin when stdin is not a terminal, so a spawned command hangs until the timeout kills it and arrives with empty stderr — a failure with no explanation, on a command that works when typed. And it assigns a MAC at runtime without recording it in device config, so MACs are derived and set explicitly, which the guest needs anyway: it names interfaces by bus position, and matching by name configures the wrong one on a multi-homed machine. No build step; Node strips the types. The lifecycle has no unit tests because a fake hypervisor would assert that the fake behaves as expected, which is the shape of test this project exists to stop shipping. |