chore(rename): tranche 2b part 2 — the podman network, migrated in three phases #369

Sammanfogat
supernaut sammanfogade 4 incheckningar från rename/tranche-2b-2-podman-network in i main 2026-08-04 12:15:56 +00:00
Ägare

ADR 0039 rename, runbook §6b-2. Renames the podman network gitborg → bitborg on both hosts.

Podman has no network rename, so this is a migration between two coexisting networks rather than
an edit. Applied and verified on both hosts; --check now reports changed=0 on each.

Why not a straight swap of Network=

The two bridges are isolated with no route between them, and containers move one role at a time.
Caddy is the front door and its role runs after everything it proxies to, so a straight swap would
502 every public endpoint from the postgres role until the caddy role — minutes, not the
per-container seconds that "recreates every container" suggests. The forgejo → caddy gap alone
measured 51 s during tranche 2a, and a flip adds postgres, kanidm and web ahead of it.

Quadlet accepts multiple Network= lines, so one flag drove three phases:

Phase Flag State Cost
1 true every container on BOTH networks 20 recreations, no partition
2 false every container on bitborg only 20 recreations, no partition
3 n/a purge the old unit + network no container touched

No phase has a partition: in phase 1 every pair still shares gitborg until both are recreated, and
in phase 2 every pair already shares bitborg. Each container is recreated twice instead of once and
the estate stays reachable throughout.

Three couplings that had to move with the name

  • The destination filename is now DERIVED from bitborg_network. It was the literal
    gitborg.network while every consumer referenced Network={{ bitborg_network }}.network, so
    flipping the variable alone would have left all 20 containers pointing at a unit that does not
    exist — and a Quadlet container whose .network is missing fails to generate, so its service
    ceases to exist rather than failing loudly.
  • The subnet is pinned, and both trusted-proxy values now derive from it.
    forgejo_reverse_proxy_trusted_proxies and caddy_trusted_proxies were the literal
    10.89.0.0/24 — the subnet podman happened to allocate, as their own comment admitted. A new
    network necessarily takes a different slot, so without this Forgejo would stop trusting Caddy's
    X-Forwarded-For, every request would resolve to Caddy's own container IP, and fail2ban would ban
    the front door (the #129/#195 shape). Confirmed live: after phase 1, DNS returned 10.89.1.x for
    every peer, so containers were already talking over the new subnet. Note the separator asymmetry —
    Forgejo takes a comma-separated list, Caddy's directive is space-separated.
  • The podman role now flushes and starts the network itself. It is the only role writing a
    Quadlet unit without meta: flush_handlers; every role writing a .container has one, which is
    why no container landed on a stale generated unit. The narrower real gap: --tags podman wrote the
    file and never created the network, because creation was otherwise a side effect of the first
    container start pulling in the generated Requires=.

The purge, and why its guard is a state predicate

Guarded on the legacy network having zero containers attached. While any .container still says
Network=gitborg.network, removing that unit file makes the next daemon-reload fail to generate that
container's service at all, orphaning a running container from systemd. The podman role runs first in
the play, so an ordering comment would be wrong on every run; a predicate cannot fire early. Verified
across all four runs: inert while 10 were attached, firing when 0 were, and each purge pass changed
exactly its three steps (stop, remove unit, podman network rm) while touching no container.

podman network rm is required and easy to miss: a Quadlet .network unit's generated ExecStart is
podman network create --ignore with no teardown, so stopping the unit and deleting its file would
leave the network and its bridge resident forever. The block is retained after the migration — it is
idempotent on a converged host and is the only thing that would clean up a host rebuilt from a
pre-rename image, matching how tranche 2a left a retirement block in every role it touched.

Two pre-existing bugs this exposed (fix(quadlet) commit)

Phase 1 reported changed=23 failed=0 while leaving kanidm and bitborg-reconcile-trigger on the
OLD network — 8 of 10 migrated. Neither defect was introduced here; neither container has a "running
container is on the pinned image" check that would have caught them earlier.

  • A Quadlet restart handler must daemon_reload itself. Handlers run in definition order
    across the play, not notification order, and the shared Reload gitborg user systemd is defined in
    several roles so Ansible de-duplicates it to the last one — runner-controller's — which sorts after
    every restart handler belonging to an earlier role. Kanidm restarted onto the stale generated unit;
    the reload then produced a correct unit nothing restarted. Both the playbook and
    systemctl --user cat kanidm.service reported the new config while the running container had the
    old one. The forgejo role had already fixed this for itself against #63; kanidm's handler carried a
    comment claiming the flush protected it, which was false. An image-tag bump would have failed
    identically.
  • state: started is a no-op on a running service. The reconcile-shim was the last container role
    without a restart-on-unit-change guard, so its unit changed and the container never moved.

daemon_reload is also folded into monitoring-agent's and runner-controller's handlers, which were
ordered correctly only because of where those roles sit in site.yml — reordering roles would have
silently stopped containers being recreated.

Operational lesson, recorded in the runbook: verify a container migration from podman ps, never
from the playbook's changed count or the generated unit.
Repairing drift that already exists needs
a one-off systemctl --user restart; once the files are correct a re-run reports changed=0 and
never notifies a restart.

Volume census

40 recreations went past the census with every named volume intact. Three shapes that look like the
ADR 0039 empty-volume failure mode in a du diff and are not, all now documented:

  • gitborg-kanidm-data moves between ~2.9M and ~5.5M across restarts — SQLite vacuum/WAL
    checkpointing. Confirm with auth.gitborg.se/status, not du.
  • gitborg-goatcounter 2.3M → 708K. Verified intact by reading the SQLite file directly:
    hit_counts and hit_stats both span 2026-06-24..2026-08-04. hits is GoatCounter's staging
    table, aggregated and cleared; that flush plus compaction is the whole drop.
  • goatcounter mints a fresh anonymous volume on every recreation — its image declares a VOLUME
    Quadlet does not map, so podman creates an unnamed volume and discards the previous one. Expected
    noise on that host, and it is what the orphan left behind by §6a actually was.

A runtime failure the syntax check could not see

Deleting the migration flag left a live when: ... not (bitborg_network_legacy_attached | bool)
referencing a variable that no longer existed, which would have failed the podman role on every run.
ansible-playbook --syntax-check passed clean because it does not evaluate when expressions; rg
for the removed variable name is what caught it.

Also fixed

§6a's verify block pointed at /data/gitea/conf/app.ini, but app.ini is bind-mounted at
/etc/gitborg/app.ini — so that documented check had been failing with "No such file or directory"
for anyone who ran it.

Out of scope

§6b-3 (nftables, admin sshd, the fail-closed user@2000 data-mount drop-in), the gitborg-* volume
units and the volumes themselves, the unix user at uid 2000, /srv/gitborg-*, the Postgres database
and role, the podman secret names, the backup archive prefix, the gitborg_* metric names, and the
host/instance labels.

ADR 0039 rename, runbook §6b-2. Renames the podman network `gitborg` → `bitborg` on both hosts. **Podman has no `network rename`**, so this is a migration between two coexisting networks rather than an edit. Applied and verified on both hosts; `--check` now reports `changed=0` on each. ## Why not a straight swap of `Network=` The two bridges are isolated with no route between them, and containers move one **role** at a time. Caddy is the front door and its role runs after everything it proxies to, so a straight swap would 502 every public endpoint from the `postgres` role until the `caddy` role — minutes, not the per-container seconds that "recreates every container" suggests. The forgejo → caddy gap alone measured 51 s during tranche 2a, and a flip adds postgres, kanidm and web ahead of it. Quadlet accepts multiple `Network=` lines, so one flag drove three phases: | Phase | Flag | State | Cost | | ----- | ------- | --------------------------------- | ---------------------------- | | 1 | `true` | every container on BOTH networks | 20 recreations, no partition | | 2 | `false` | every container on `bitborg` only | 20 recreations, no partition | | 3 | n/a | purge the old unit + network | no container touched | No phase has a partition: in phase 1 every pair still shares `gitborg` until both are recreated, and in phase 2 every pair already shares `bitborg`. Each container is recreated twice instead of once and the estate stays reachable throughout. ## Three couplings that had to move with the name - **The destination filename is now DERIVED from `bitborg_network`.** It was the literal `gitborg.network` while every consumer referenced `Network={{ bitborg_network }}.network`, so flipping the variable alone would have left all 20 containers pointing at a unit that does not exist — and a Quadlet container whose `.network` is missing fails to **generate**, so its service ceases to exist rather than failing loudly. - **The subnet is pinned, and both trusted-proxy values now derive from it.** `forgejo_reverse_proxy_trusted_proxies` and `caddy_trusted_proxies` were the literal `10.89.0.0/24` — the subnet podman happened to allocate, as their own comment admitted. A new network necessarily takes a different slot, so without this Forgejo would stop trusting Caddy's `X-Forwarded-For`, every request would resolve to Caddy's own container IP, and fail2ban would ban the front door (the #129/#195 shape). Confirmed live: after phase 1, DNS returned `10.89.1.x` for every peer, so containers were already talking over the new subnet. Note the separator asymmetry — Forgejo takes a comma-separated list, Caddy's directive is space-separated. - **The podman role now flushes and starts the network itself.** It is the only role writing a Quadlet unit without `meta: flush_handlers`; every role writing a `.container` has one, which is why no container landed on a stale generated unit. The narrower real gap: `--tags podman` wrote the file and never created the network, because creation was otherwise a side effect of the first container start pulling in the generated `Requires=`. ## The purge, and why its guard is a state predicate Guarded on the legacy network having **zero containers attached**. While any `.container` still says `Network=gitborg.network`, removing that unit file makes the next daemon-reload fail to generate that container's service at all, orphaning a running container from systemd. The podman role runs first in the play, so an ordering comment would be wrong on every run; a predicate cannot fire early. Verified across all four runs: inert while 10 were attached, firing when 0 were, and each purge pass changed exactly its three steps (stop, remove unit, `podman network rm`) while touching no container. `podman network rm` is required and easy to miss: a Quadlet `.network` unit's generated `ExecStart` is `podman network create --ignore` with no teardown, so stopping the unit and deleting its file would leave the network and its bridge resident forever. The block is retained after the migration — it is idempotent on a converged host and is the only thing that would clean up a host rebuilt from a pre-rename image, matching how tranche 2a left a retirement block in every role it touched. ## Two pre-existing bugs this exposed (`fix(quadlet)` commit) Phase 1 reported `changed=23 failed=0` while leaving `kanidm` and `bitborg-reconcile-trigger` on the OLD network — 8 of 10 migrated. Neither defect was introduced here; neither container has a "running container is on the pinned image" check that would have caught them earlier. - **A Quadlet restart handler must `daemon_reload` itself.** Handlers run in **definition** order across the play, not notification order, and the shared `Reload gitborg user systemd` is defined in several roles so Ansible de-duplicates it to the last one — runner-controller's — which sorts after every restart handler belonging to an earlier role. Kanidm restarted onto the stale generated unit; the reload then produced a correct unit nothing restarted. Both the playbook **and** `systemctl --user cat kanidm.service` reported the new config while the running container had the old one. The forgejo role had already fixed this for itself against #63; kanidm's handler carried a comment claiming the flush protected it, which was false. An image-tag bump would have failed identically. - **`state: started` is a no-op on a running service.** The reconcile-shim was the last container role without a restart-on-unit-change guard, so its unit changed and the container never moved. `daemon_reload` is also folded into monitoring-agent's and runner-controller's handlers, which were ordered correctly only because of where those roles sit in `site.yml` — reordering roles would have silently stopped containers being recreated. **Operational lesson, recorded in the runbook: verify a container migration from `podman ps`, never from the playbook's `changed` count or the generated unit.** Repairing drift that already exists needs a one-off `systemctl --user restart`; once the files are correct a re-run reports `changed=0` and never notifies a restart. ## Volume census 40 recreations went past the census with every named volume intact. Three shapes that look like the ADR 0039 empty-volume failure mode in a `du` diff and are not, all now documented: - `gitborg-kanidm-data` moves between ~2.9M and ~5.5M across restarts — SQLite vacuum/WAL checkpointing. Confirm with `auth.gitborg.se/status`, not `du`. - `gitborg-goatcounter` 2.3M → 708K. Verified intact by reading the SQLite file directly: `hit_counts` and `hit_stats` both span 2026-06-24..2026-08-04. `hits` is GoatCounter's staging table, aggregated and cleared; that flush plus compaction is the whole drop. - goatcounter mints a fresh **anonymous** volume on every recreation — its image declares a `VOLUME` Quadlet does not map, so podman creates an unnamed volume and discards the previous one. Expected noise on that host, and it is what the orphan left behind by §6a actually was. ## A runtime failure the syntax check could not see Deleting the migration flag left a live `when: ... not (bitborg_network_legacy_attached | bool)` referencing a variable that no longer existed, which would have failed the podman role on every run. `ansible-playbook --syntax-check` passed clean because it does not evaluate `when` expressions; `rg` for the removed variable name is what caught it. ## Also fixed §6a's verify block pointed at `/data/gitea/conf/app.ini`, but app.ini is bind-mounted at `/etc/gitborg/app.ini` — so that documented check had been failing with "No such file or directory" for anyone who ran it. ## Out of scope §6b-3 (nftables, admin sshd, the fail-closed `user@2000` data-mount drop-in), the `gitborg-*` volume units and the volumes themselves, the unix user at uid 2000, `/srv/gitborg-*`, the Postgres database and role, the podman secret names, the backup archive prefix, the `gitborg_*` metric names, and the host/instance labels.
supernaut lade till 4 incheckningar 2026-08-04 12:03:23 +00:00
ADR 0039 rename, runbook §6b-2, phase 1 of 3. Renames the podman network by MIGRATING between two
coexisting networks rather than swapping the name, because podman has no `network rename`.

### Why not a straight swap

The two bridges are isolated with no route between them, and containers move one ROLE at a time.
Caddy is the front door and its role runs after everything it proxies to, so a straight swap of
`Network=` would 502 every public endpoint from the `postgres` role until the `caddy` role — minutes,
not the per-container seconds that "recreates every container" suggests. The forgejo -> caddy gap
alone measured 51s during tranche 2a, and a flip adds postgres, kanidm and web ahead of it.

Quadlet accepts multiple `Network=` lines, so one flag drives three phases:

  phase 1  bitborg_network_legacy_attached: true   -> every container on BOTH networks
  phase 2  bitborg_network_legacy_attached: false  -> every container on `bitborg` only
  phase 3  the podman role's guarded purge collects the old unit + podman network

No phase has a partition: in phase 1 every pair still shares `gitborg` until both are recreated, and
in phase 2 every pair already shares `bitborg`. The cost is two recreations per container instead of
one, and the estate stays reachable throughout.

### Three couplings that had to move with the name

- **The destination filename is now DERIVED from `bitborg_network`.** It was the literal
  `gitborg.network` while every consumer referenced `Network={{ bitborg_network }}.network`, so
  flipping the variable alone would have left all 20 containers referencing a unit that does not
  exist — and a Quadlet container whose `.network` is missing fails to GENERATE, so its service
  ceases to exist rather than failing loudly.
- **The subnet is pinned, and both trusted-proxy values now derive from it.**
  `forgejo_reverse_proxy_trusted_proxies` and `caddy_trusted_proxies` were the literal
  `10.89.0.0/24` — the subnet podman happened to allocate, as their own comment admitted. A new
  network necessarily takes a different slot, so without this Forgejo would stop trusting Caddy's
  X-Forwarded-For, every request would resolve to Caddy's own container IP, and fail2ban would ban
  the front door (the #129/#195 shape). Pinning the subnet is what makes the value stateable in
  advance. Note the separator asymmetry: Forgejo takes a comma-separated list, Caddy's directive is
  space-separated.
- **The podman role now flushes and starts the network itself.** It is the only role writing a
  Quadlet unit without `meta: flush_handlers`; every role writing a `.container` has one, which is
  why the reload does happen and no container lands on a stale generated unit. The narrower real gap
  is that `--tags podman` wrote the file and never created the network, because creation was
  otherwise a side effect of the first container start pulling in the generated `Requires=`.

### The purge, and why its guard is a predicate

Included now but inert until phase 2, guarded on the legacy network having ZERO containers attached.
While any `.container` still says `Network=gitborg.network`, removing that unit file makes the next
daemon-reload fail to generate that container's service at all, orphaning a running container from
systemd. The podman role runs first in the play, so an ordering comment would be wrong on every run;
a state predicate cannot fire early. It also runs `podman network rm`, which is easy to miss: a
Quadlet `.network` unit's generated ExecStart is `podman network create --ignore` with no teardown,
so stopping the unit and deleting its file leaves the network and its bridge resident forever.

### Dry-run

prod `changed=20 failed=0`, monitoring `changed=3 failed=0` (looped tasks count once: the network
unit, 10 container units, 10 restarts). Every changed task is a `.container` gaining its second
`Network=` line, that container's restart, or one of the two trusted-proxy files. On both hosts the
legacy network unit renders **`ok`, not `changed`** — the parameterised template reproduces the live
file byte for byte rather than churning a network podman would refuse to re-subnet anyway.

`bitborg-reconciler` needs no change: it is `Network=host` (ADR 0017), which is also why the
attached count is 10+10 rather than 11+10.
Two pre-existing latent bugs, both exposed by ADR 0039 tranche 2b-2 phase 1 and neither introduced
by it: the play reported `changed=23 failed=0` while leaving `kanidm` and `bitborg-reconcile-trigger`
running on the OLD podman network, with 8 of 10 containers migrated. A converged playbook did not
mean the running container matched its own unit file.

### Handlers run in DEFINITION order, and the shared reload sorts last

`Reload gitborg user systemd` is defined in several roles, so Ansible de-duplicates it to the LAST
definition — the runner-controller role's — which places it after every restart handler belonging to
an earlier role. At the kanidm flush the observed order was:

    RUNNING HANDLER [kanidm : Restart Kanidm]                          <- first
    RUNNING HANDLER [runner-controller : Reload gitborg user systemd]  <- too late

Quadlet is a generator, so a unit is regenerated only at daemon-reload. Kanidm therefore restarted
onto the stale generated unit; the reload then produced a correct unit that nothing restarted. The
failure is silent in both directions — `systemctl --user cat kanidm.service` shows the NEW
`--network bitborg --network gitborg`, while the running container has only the old one.

The forgejo role already folds `daemon_reload` into its own restart handler and records #63 as the
reason. The kanidm handler's comment claimed "flush_handlers in the role ensures the unit exists
(daemon-reload) before this runs", which was simply false: the flush runs the handlers, and this
handler went first. An image-tag bump would have failed the same way, and unlike
forgejo/postgres/caddy/web this role has no "running container is on the pinned image" check to catch
it (worth a follow-up).

Fixed by folding `daemon_reload: true` into the restart handler, exactly as forgejo does.

### The shim never restarted on a unit change

`state: started` is a no-op on a running service, so a re-rendered unit could never come live — the
#63 shape again. Every other container role already guards this (postgres, web and caddy inline;
forgejo and now kanidm via the handler). The shim was the last one missing it, so its unit gained a
second `Network=` line, the play reported changed, and the container kept running on the old network.
An image bump would have no-opped identically.

### Hardening: make the invariant local instead of global

monitoring-agent's three handlers and runner-controller's were ordered correctly ONLY because of
where those roles sit in site.yml — monitoring-agent after runner-controller, and
runner-controller's own Reload defined above its Restart in the same file. Reordering site.yml would
have silently stopped containers being recreated, with no error and nothing to catch it. Folding
`daemon_reload` into each makes them correct on their own terms. A redundant reload is a cheap no-op
on a user manager this size.

Note what this does NOT fix: the two containers already running stale on prod. Their unit files and
generated units are correct, so a re-run reports `changed=0` and never notifies a restart. Repairing
existing drift needs a one-off `systemctl --user restart` of each; this commit only stops it
recurring.
ADR 0039 rename, runbook §6b-2, phase 2 of 3. Flips `bitborg_network_legacy_attached` to false, which
in one variable: removes the second `Network=` line from all 20 .container templates, drops the legacy
entry from `podman_networks` so the role stops managing `gitborg.network`, and narrows both
trusted-proxy values to `10.89.1.0/24` alone.

Partition-free for the mirror of phase 1's reason: a container still holding both networks can always
reach one already narrowed to `bitborg`, so the estate stays connected while containers move a role at
a time. Downtime is per-container restart only.

Phase 1 verified first, from `podman ps` rather than from the playbook's `changed` count — the count
was green on prod while two containers had not moved. Both hosts: 10/10 on both networks, volume
census unchanged, 6/6 probes green, both node-exporters reporting (so prod -> monitoring remote_write
survived both hosts' network changes).

### Dry-run

prod `changed=19 failed=0`, monitoring `changed=2 failed=0` (looped tasks count once: 10 unit
rewrites, 10 restarts). Every changed task is a `.container` losing its transitional line, that
container's restart, or one of the two trusted-proxy files. Three things confirm the mechanism:

- `Restart Kanidm` now reports changed, where the unfixed handler silently restarted onto a stale
  generated unit;
- the network task manages only `item=bitborg` and reports `ok`;
- the purge tasks report `skipping`, which is the documented `not ansible_check_mode` limitation, NOT
  the predicate. On the real phase-2 run the predicate finds 10 containers still attached and stays
  inert; the purge only fires on the phase-3 pass, when the count reaches zero.

### Two volume movements chased down during phase 1, both benign

Recorded because 20 recreations is exactly where the ADR 0039 empty-volume failure mode lives, and
both look alarming in a `du` diff:

- `gitborg-kanidm-data` 5.5M -> 2.9M: SQLite vacuum on each start; `auth.gitborg.se/status` green
  throughout.
- `gitborg-goatcounter` 2.3M -> 708K: read the SQLite file directly — `hit_counts` and `hit_stats`
  both span 2026-06-24..2026-08-04, 30 paths, 1 site. Six weeks of history intact. `hits` is
  GoatCounter's staging table, aggregated into the count tables and cleared; that flush plus
  compaction is the whole drop.

Also worth knowing before the next recreation: goatcounter's image declares a `VOLUME` at
`/home/goatcounter/goatcounter-data` that Quadlet does not map, so podman mints a fresh ANONYMOUS
volume on every recreation and discards the previous one. It is empty and unused (the real data is
the named volume at `/data`). A new anonymous volume after an apply on that host is expected noise,
not a signal — it is what the orphan left behind by tranche 2a actually was.
ADR 0039 rename, runbook §6b-2, phase 3 of 3. The purge ran on both hosts and this removes the
migration scaffolding: the `bitborg_network_legacy_attached` flag, the two legacy vars, and the
`{% if %}` block from all 20 .container templates.

### Result on both hosts

`bitborg` at the pinned 10.89.1.0/24 with all 10 containers attached to it alone; the `gitborg`
network and its `.network` unit gone; `--check` reporting `changed=0`. Volume census unchanged
throughout. Each purge pass changed exactly its three steps (stop, remove unit, `podman network rm`)
and touched no container — the state predicate was inert while 10 were attached and fired when 0
were, across all four runs.

The scaffolding removal renders byte-identically to the flag-false state, verified rather than
assumed: `changed=0` on prod AND monitoring after the edit.

### The purge block is deliberately RETAINED

It is idempotent on a converged host — the network is gone so the attachment check returns empty, the
unit file is absent and both removals no-op — and it is the only thing that would clean up a host
rebuilt or restored from a pre-rename image. Same reasoning as tranche 2a leaving a retirement block
in every role it touched.

### A runtime failure the syntax check could not see

Deleting the flag left a live `when: ... not (bitborg_network_legacy_attached | bool)` on the
attachment-check task, referencing a variable that no longer existed — which would have failed the
podman role on every run. `ansible-playbook --syntax-check` passed clean, because it does not
evaluate `when` expressions; `rg` for the removed variable name is what caught it. Worth remembering
that the cheap validation and the meaningful one are not the same gate.
supernaut sammanfogade incheckning 0c53bf079b till main 2026-08-04 12:15:56 +00:00
supernaut tog bort grenen rename/tranche-2b-2-podman-network 2026-08-04 12:15:56 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!369
Ingen beskrivning angiven.