The bed's 058 claim, base64 dump, and 063 comment are made honest (review)
- The RestartCount==0 assertion cannot discriminate the 058 fix: the consumer is built after its broker is already up, so patient and exit-on-unreachable code both connect first-try; and restart-on recreates reset the count. Downgraded to an honest liveness check and the comment now points at the mesh-tools unit test as the real proof. - The trust-failure dump ran base64 -d over declared.json, which is JSON (not base64), so it always reported 'no trust' — removed; the adjacent python check that decodes the inner declaration field is kept. - The 063 comment claimed the vhost is re-listed after the restart; the code only runs a TCP probe. Comment corrected to what the code proves, and the docker-proxy-vs-DNAT coupling is noted.
This commit is contained in:
@@ -258,7 +258,10 @@ test("a joined node's consumers open the store and broker the mesh built and ado
|
||||
`\n--- ${machine} mesh-host.log ---\n` +
|
||||
(await on(machine, `tail -60 /var/log/mesh-host.log 2>&1`)).out +
|
||||
`\n--- ${machine} declared registry-trust? ---\n` +
|
||||
(await on(machine, `base64 -d < /var/lib/mesh-host/declared.json 2>/dev/null | grep -c registry-trust; python3 -c "import json,base64; d=json.load(open('/var/lib/mesh-host/declared.json')); print('registry-trust' in base64.b64decode(d['declaration']).decode())" 2>&1`)).out);
|
||||
// declared.json is {"declaration":"<base64>","signature":"<base64>"} — JSON, not base64
|
||||
// itself, so only the inner `declaration` field is decoded. (An earlier `base64 -d` on the
|
||||
// whole file always errored and reported "no trust" whether or not the trust was declared.)
|
||||
(await on(machine, `python3 -c "import json,base64; d=json.load(open('/var/lib/mesh-host/declared.json')); print('registry-trust' in base64.b64decode(d['declaration']).decode())" 2>&1`)).out);
|
||||
}
|
||||
|
||||
// The shared base first — every module with code of its own stands on it.
|
||||
@@ -317,26 +320,33 @@ test("a joined node's consumers open the store and broker the mesh built and ado
|
||||
for (const c of ["mesh-store", "mesh-broker", "mesh-controller"]) {
|
||||
assert.match(up, new RegExp(`(^|\\n)${c}(\\n|$)`), `${c} did not survive adoption:\n${up}`);
|
||||
}
|
||||
// Steady a moment, then confirm the consumer did not crash-loop after connecting.
|
||||
// Steady a moment, then confirm the consumer is up and did not crash-loop after connecting.
|
||||
//
|
||||
// **This is a liveness check, not the proof of hq issue 058.** By the time amqp-ping is built
|
||||
// and assigned, its broker is already up and reachable, so both the patient runtime and the
|
||||
// old exit-on-unreachable code connect on the first try — a green here does not discriminate
|
||||
// the fix. And RestartCount counts in-place restarts of the *current* container, which a
|
||||
// restart-on recreate (amqp-ping declares `restart-on: [amqp-env]`) resets to zero. The honest
|
||||
// proof of the patient-reconnect logic is the mesh-tools unit test of fatalBrokerReason; what
|
||||
// this asserts is only that the consumer settled and stayed up.
|
||||
await new Promise((r) => setTimeout(r, 15_000));
|
||||
assert.match((await on(NODE, `docker ps --format '{{.Names}}\t{{.Status}}'`)).out,
|
||||
/amqp-ping\tUp/, `amqp-ping did not stay up on ${NODE}`);
|
||||
// And was never restarted by the container runtime at all (hq issue 058): a runtime whose
|
||||
// broker is not reachable yet waits for it in-process, so "the overlay came up a moment after
|
||||
// the container" must produce zero restarts — churn here is the crash-loop 058 retired.
|
||||
const restarts = (await on(NODE, `docker inspect -f '{{.RestartCount}}' amqp-ping`)).out.trim();
|
||||
assert.equal(restarts, "0",
|
||||
`amqp-ping was restarted ${restarts} time(s) by the container runtime on ${NODE}; ` +
|
||||
`a runtime waits for its broker in-process (hq issue 058):\n` +
|
||||
(await on(NODE, `docker logs amqp-ping 2>&1 | head -20`)).out);
|
||||
/amqp-ping\tUp/, `amqp-ping did not stay up on ${NODE}:\n` +
|
||||
(await on(NODE, `docker logs amqp-ping 2>&1 | tail -20`)).out);
|
||||
|
||||
// The broker survives a restart as a reachable thing, not just a running one (hq issue 063).
|
||||
// The foundation's ports are opened from anywhere on input, but the broker is a published
|
||||
// container port — so a cross-node dial is forwarded, and until 063 the forward chain carried
|
||||
// no foundation rule: the joined node reached the broker only until its first connection's
|
||||
// conntrack entry dropped. Restarting the broker here drops it deliberately; the node must
|
||||
// reconnect and the vhost the provisioner holds must still be listed, proving the forward rule
|
||||
// and not a surviving conntrack entry is what carries the connection.
|
||||
// conntrack entry dropped. Restarting the broker here drops it deliberately, then a FRESH TCP
|
||||
// connection from the joined node must complete — a new 5-tuple in state NEW cannot ride the
|
||||
// old established entry, so only the forward rule 063 adds can carry it. (This asserts the
|
||||
// reachability the fix restores; the vhost minting itself is already asserted above.)
|
||||
//
|
||||
// The probe reaches the forward chain because this runtime DNATs published ports (no userland
|
||||
// proxy) — a cross-node dial is redirected to the container and forwarded. Were docker-proxy
|
||||
// ever enabled, the same dial would terminate on the host and be satisfied by the INPUT rule,
|
||||
// masking a missing forward rule; the lab's runtime does not use it, so the probe is honest here.
|
||||
await must(CONTROL, `docker restart mesh-broker`, 120_000);
|
||||
{
|
||||
const deadline = Date.now() + 180_000;
|
||||
|
||||
Reference in New Issue
Block a user