novox/hq ADR 0035: one implementation, several surfaces, and a surface
holds no decisions. The act of assigning — including that an assignment
which does not resolve is kept and still refused — moved into acts.go,
and the command line now calls it too. Two surfaces, one refusal, in the
same words.
It will not run without --issuer, and refuses at start rather than per
request so it is found by whoever ran it rather than by whoever finds
it. There is no flag that removes the check.
The authenticator is honest about what it is: no token can be verified
until an identity provider exists, because that is a module and none is
running, so every request is refused and told that the command line
still works. A surface that functioned without authentication would be
one somebody left running — and the board this stands behind is
published on a public name.
Four refusals, four tests. The last one first asserted "not 200", which
passed because a request with no database fails at the store anyway — it
proved nothing about whether the input was checked. It now asserts the
specific refusal, and bites when the check is removed.
Manifests for a provider and a consumer, so the contract can be read
rather than only exercised through a lab fixture that stages the grants
by hand.
Checked as a pair rather than separately, because two manifests that
only ever parse alone are two manifests nobody has held against each
other. The test asserts the names match, that each side says where it
wants to be told, and that the consumer contributes the key the
provisioner actually reads.
That last one is the trap worth having a test for: a consumer
contributing "name" — which is exactly what a database consumer
contributes — resolves cleanly, deploys, and then fails on the machine
with "asked for a bucket and did not name it". Nothing in that message
points back at the manifest that caused it. Both mistakes were made
while writing these two files.
dnsmasq read /etc/resolv.conf to find where to forward. Whatever points a
machine at the mesh writes its own address into that file — so dnsmasq's
upstream was dnsmasq, and every query it could not answer locally looped. Its
receive queue filled with 15KB of them and every lookup on the machine hung,
which is why this arrived as a thirty-second timeout rather than a wrong
answer.
It needs no upstream at all: the asking module routes only the mesh's suffix
here and leaves everything else where the machine already sent it. And it names
none, because choosing one would send every query this machine makes somewhere
nobody agreed to.
Also corrected: the comment claiming it takes only 127.0.0.55. Listening on a
loopback address makes dnsmasq take the rest of loopback with it, 127.0.0.1
included — which is what claiming `the-dns-port` already says, and which the
comment was quietly denying. That is the same comfortable claim as ".54 is
free", in the same file, made twice.
`127.0.0.54` is systemd-resolved's DNS *proxy* stub. The module asserted it was
free, in a comment that read as reasoned — "not .53, that is
systemd-resolved's" — and it was simply wrong: resolved holds both. dnsmasq
could not create the socket and never started.
Nothing in a unit test could have caught it. They checked the module names an
address and that the asking modules point at the same one, and all of that
passed while the daemon could not start. Only a machine knows which addresses
are spare, which is the argument for proving a module that asserts facts about
machines on a machine, before believing the assertions.
So it moves to .55, and says what that is: a convention, not a reservation. If
a future systemd takes it, this line changes and nothing else does.
The tests now derive the address from the serving module and check the two
asking modules agree with it, rather than naming it a fourth time — that fourth
place is the one nobody would think to change.
And the lab assigns `resolved-split-dns` rather than `resolv-conf`: those
machines run systemd-resolved, which owns the file. The two claim the same
thing precisely so the wrong choice is a refusal rather than a fight, and
picking the wrong one was testing the fight.
Three manifests and the rule that keeps them apart. Serving and asking are
genuinely different roles, and systemd-resolved can only do the second — it
cannot answer a wildcard, it routes the mesh's suffix to something that can. A
module that treated them as one role could not work, which is the mistake worth
naming rather than discovering.
So `the-dns-port` and `the-resolver-configuration` are two claims. A machine
gets one of each, and two of either is refused by the mesh rather than fought
over on the machine — which is what ADR 0009's table meant by listing resolvers
beside the seat and pid 1. That table names the resource `/etc/resolv.conf`,
which is what it is; a claim is a name in the catalogue's own form, and the
catalogue refuses the path as one.
Neither module knows anything about the machine it is on, which is what lets
them be static manifests: they name `mesh0` and `127.0.0.54`, both chosen by
the mesh, rather than an address only that machine has. Not 127.0.0.1 and not
127.0.0.53 — taking either would be a module claiming something it did not say
it claims.
A service can now reflect a file another module put on the machine, written
`<module>.<id>`. The resolver has to restart when the mesh rewrites the names;
without it, it would serve the names it started with for ever, with every
machine that joined afterwards unreachable and every check passing.