papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist.
79 lines
3.5 KiB
Markdown
79 lines
3.5 KiB
Markdown
---
|
|
status: accepted
|
|
date: 2026-08-04
|
|
deciders: jochen
|
|
reconstructed: true
|
|
---
|
|
|
|
# 14. Build, publish and deploy are three silos with different cardinality
|
|
|
|
> Reconstructed after the fact from the evidence cited below.
|
|
|
|
## Context
|
|
|
|
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
|
|
run a different number of times.
|
|
|
|
- Compiling happens **once per module feature**, on the build node.
|
|
- Packaging and uploading happens **once per module feature**, on the build node.
|
|
- Installing, configuring, starting and verifying happens **once per module feature per node**.
|
|
|
|
Conflating them is what made earlier versions slow and hard to reason about. Work that should
|
|
happen once was being repeated per node, and the fan-out point was implicit rather than a
|
|
boundary anything could observe.
|
|
|
|
The split had been declared before it was real. Packaging still happened inside the build,
|
|
which meant the boundary existed in the documentation and not in the code.
|
|
|
|
## Considered options
|
|
|
|
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
|
|
Cardinality is then a property of each stage's implementation, and nothing can reason about
|
|
the pipeline as a whole.
|
|
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
|
|
so build must know every module, every feature, and how each composes its artifact —
|
|
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
|
|
a stale package instead of re-packaging.
|
|
3. **Three silos, with an explicit handover between each.** Chosen.
|
|
|
|
## Decision
|
|
|
|
Delivery is three silos, and the boundaries are real:
|
|
|
|
| Silo | Runs | Where |
|
|
|---|---|---|
|
|
| **build** | once per module feature | the build node |
|
|
| **publish** | once per module feature | the build node |
|
|
| **deploy** | once per module feature **per node** | every assigned node |
|
|
|
|
Commands and events are addressed **per feature**, not per module.
|
|
|
|
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
|
|
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
|
|
publishing, so a module whose artifact is a package publishes in the publish silo, not the
|
|
build one.
|
|
|
|
Modules are resolved into dependency **levels**, and a level completes before the next begins,
|
|
so a module always builds against its dependencies' freshly published versions.
|
|
|
|
## Consequences
|
|
|
|
- Work that should happen once happens once. The fan-out point is explicit and observable.
|
|
- A failed upload retries by re-packaging, because packaging belongs to the stage that
|
|
uploads.
|
|
- The handover is a staged tree in a known location rather than the build's working directory,
|
|
which is reference-counted and cannot be assumed to still exist when a later stage runs.
|
|
- The build node is now the only node that has already passed through two silos when the
|
|
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
|
|
stages; code that knew only about the first parked the build node forever while every other
|
|
node deployed cleanly.
|
|
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
|
|
reports success over a stall it cannot see.
|
|
|
|
## References
|
|
|
|
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
|
|
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
|
|
documents claimed otherwise and were stale until 2026-08-06.
|
|
- The build-node stage-tracking failure was observed on pipeline #5557.
|