Tier 0: full-stack smoke test (Molecule) — assert every service container starts #104
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#104
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Tier 0 — the cheapest guard, and the one that would have caught #94. A CI job that converges the container roles and asserts each service actually starts and responds, not just that the templates render.
*.serviceisactive;podman psshows the container up (not crash-looping); HTTP smoke on caddy (all host blocks answer), forgejo/api/healthz, etc.NoNewPrivileges=true+DropCapability=ALL) fronting an image whose binary carries file capabilities must stillexec— a regression here must fail CI.ansible/.Note the parity gap (rootless podman in CI ≠ Debian host for sysctls/pasta) — that's what Tier 1 covers; Tier 0 catches container-start/config failures for free.
Epic: gitborg/gitborg-docs#37
Increment 1 shipped — PR #110 (container-start smoke)
scripts/smoke-containers.py+ a CI step: for each hardened Quadlet unit (NoNewPrivileges+DropCapability=ALL) with a public image,podman runit with the same caps and assert the entrypoint can exec — catching the #81/#94 class (file-cap binary crash-looping when a needed cap is dropped) that--checkcan't see. Runs on the host-backend runner (ADR 0021). Validated both directions locally: passes on the fixed units; fails when the #94 regression is re-introduced.Feasibility note that shaped the scope: a full multi-service Molecule converge (as this issue originally sketched) needs systemd + rootless podman + podman-in-podman + the registry image + source-built Kanidm + the ADR-0025 fail-closed volume — flaky and costly for an every-PR tier, and it belongs on a real host. So Tier-0 = focused behavioural smoke of the actual failure classes; the full converge is Tier-1 (#105) on the ephemeral VM.
Remaining Tier-0 scope
.prom-writing script (or a factored core) and assert the output is world-readable (0644) and passespromtool check metrics. Thetoken-audit.prom0600 bug node_exporter silently rejected would be caught here.Full multi-service converge stays out of Tier-0 by design → #105.
Increment 2 — PR #111 (textfile-metric guard)
The remaining #96-class item is done:
scripts/smoke-textfile.pyfails any metric script that publishes its.promfrom amktemp(0600) file without a world-readablechmod— traced to the actualmvsource, so it's precise (no false positives on the 9 current scripts; teeth confirmed by reverting token-audit's chmod). Chosen static over behavioural on purpose — the faithfulnode_textfile_scrape_error==0check needs a real converge and belongs in Tier-1 / the post-apply gate (#105 / #107).Tier-0 status: complete
Both prod incident classes this week are now guarded in CI:
Full multi-service integration converge remains out of Tier-0 by design → Tier-1 (#105). So this issue can close once #111 merges; #105 carries the converge.
Related: PR #112 gates all CI steps by changed area (docs-only PRs skip the infra toolchain) with the directional ansible↔opentofu cross-dependency handled — separate hygiene improvement, stacked on #111.
Tier-0 is complete and in
main:Both prod incident classes from this week are now guarded on every applicable PR, and CI only runs the infra toolchain when the relevant area changed (#112). The full multi-service integration converge stays out of Tier-0 by design → Tier-1 (#105). Closing; #105 carries the remainder.