fix: bound log retention to 7 days and cover every table in user erasure #532

Öppen
supernaut vill sammanfoga 6 incheckningar från s[2]s in i main
Ägare

What

  1. base: journald gets MaxRetentionSec=7day and MaxFileSec=1day on both hosts (variables base_journal_max_retention, base_journal_max_file_sec).
  2. monitoring: no config change needed beyond 1. The Caddy comment was wrong and is fixed.
  3. Runbook: operational auth/git/www/grafana/stats/ntfy.gitborg.se names in commands, URLs and the Grafana admin group SPN (grafana_admins@auth.bitborg.se, matches kanidm_domain and grafana_admin_group) now use bitborg.se. Switchover history, redirect descriptions, past measurements and example email addresses are left as written.
  4. Runbook user-erasure section covers every personal-data table and uses auth.bitborg.se.

Why

The privacy policy says logs are kept at most 7 days. journald had a 500 MB cap and no time limit. MaxFileSec=1day is needed because journald only vacuums archived files and rotates monthly by default, so MaxRetentionSec alone would not bite.

Findings

  • Monitoring VM Caddy writes its access log (client IPs) to stdout. The podman log driver is journald, so it lands in that host's journald. There is no Alloy on the monitoring host, so it never reached Loki (the old comment said it did). The journald limit in 1 now bounds it to 7 days. No Caddy change.
  • Services host Caddy writes /var/log/caddy/access.log with roll_size 20MiB, roll_keep 5, roll_keep_for 168h. Already set, no change. Loki retention is 168h already.
  • Ceiling on the services host file: Caddy has no time-based rotation. It deletes aged rolled files only when it rotates, and never touches the active file. On a quiet host the active file could hold lines older than 7 days. Not changed (no deterministic fix without a timer or a smaller roll_size). Needs the actual file age on prod to judge.
  • Erasure: email_change_requests (holds the new address, keyed by Kanidm sub) is now deleted. Step 1 reads the sub before the Kanidm delete. Kept on purpose: checkout_consents, withdrawal_requests (sub only, no name or email) and the billing database (accounting, 7 years).

Check-mode output (--check --diff --tags base, hosts bitborg-prod and bitborg-monitoring)

--- before: /etc/systemd/journald.conf.d/10-bitborg-persistent.conf
+++ after: /etc/systemd/journald.conf.d/10-bitborg-persistent.conf
 SystemMaxUse=500M
 RuntimeMaxUse=100M
+MaxRetentionSec=7day
+MaxFileSec=1day

Recap: bitborg-prod ok=29 changed=2, bitborg-monitoring ok=27 changed=2 (the drop-in and the Restart journald handler).

--tags monitoring on bitborg-monitoring: the Caddyfile comment change renders (1 task changed) and the Caddy front-door restart handler fires. Comment only, no behaviour change.

Apply (operator)

cd ansible
ansible-playbook site.yml --limit bitborg-prod,bitborg-monitoring --tags base
ansible-playbook site.yml --limit bitborg-monitoring --tags monitoring

Then journalctl --disk-usage and journalctl --header | grep -i "head realtime" after a day to confirm rotation.

Tested

  • ansible-lint: 0 failures, 0 warnings, 57 files.
  • pnpm format:check, pnpm mdlint: clean. Pre-commit hooks passed on all three commits.
  • Check mode as above. No apply.

Open questions

  • Services host Caddy active file age (see ceiling above).
  • Billing database: whether a request to de-identify payment records is ever honoured is an operator decision. The runbook says this procedure does not touch it.
  • Legal basis for keeping checkout_consents and withdrawal_requests is my reading (consumer-law evidence, accounting trail). Confirm.
  • Other auth.gitborg.se mentions remain elsewhere in the runbook (bootstrap, Kanidm and monitoring sections). Out of scope here.
## What 1. `base`: journald gets `MaxRetentionSec=7day` and `MaxFileSec=1day` on both hosts (variables `base_journal_max_retention`, `base_journal_max_file_sec`). 2. `monitoring`: no config change needed beyond 1. The Caddy comment was wrong and is fixed. 3. Runbook: operational `auth/git/www/grafana/stats/ntfy.gitborg.se` names in commands, URLs and the Grafana admin group SPN (`grafana_admins@auth.bitborg.se`, matches `kanidm_domain` and `grafana_admin_group`) now use `bitborg.se`. Switchover history, redirect descriptions, past measurements and example email addresses are left as written. 4. Runbook user-erasure section covers every personal-data table and uses `auth.bitborg.se`. ## Why The privacy policy says logs are kept at most 7 days. journald had a 500 MB cap and no time limit. `MaxFileSec=1day` is needed because journald only vacuums archived files and rotates monthly by default, so `MaxRetentionSec` alone would not bite. ## Findings - Monitoring VM Caddy writes its access log (client IPs) to stdout. The podman log driver is journald, so it lands in that host's journald. There is no Alloy on the monitoring host, so it never reached Loki (the old comment said it did). The journald limit in 1 now bounds it to 7 days. No Caddy change. - Services host Caddy writes `/var/log/caddy/access.log` with `roll_size 20MiB`, `roll_keep 5`, `roll_keep_for 168h`. Already set, no change. Loki retention is 168h already. - Ceiling on the services host file: Caddy has no time-based rotation. It deletes aged rolled files only when it rotates, and never touches the active file. On a quiet host the active file could hold lines older than 7 days. Not changed (no deterministic fix without a timer or a smaller `roll_size`). Needs the actual file age on prod to judge. - Erasure: `email_change_requests` (holds the new address, keyed by Kanidm `sub`) is now deleted. Step 1 reads the `sub` before the Kanidm delete. Kept on purpose: `checkout_consents`, `withdrawal_requests` (`sub` only, no name or email) and the billing database (accounting, 7 years). ## Check-mode output (`--check --diff --tags base`, hosts bitborg-prod and bitborg-monitoring) ``` --- before: /etc/systemd/journald.conf.d/10-bitborg-persistent.conf +++ after: /etc/systemd/journald.conf.d/10-bitborg-persistent.conf SystemMaxUse=500M RuntimeMaxUse=100M +MaxRetentionSec=7day +MaxFileSec=1day ``` Recap: bitborg-prod ok=29 changed=2, bitborg-monitoring ok=27 changed=2 (the drop-in and the `Restart journald` handler). `--tags monitoring` on bitborg-monitoring: the Caddyfile comment change renders (1 task changed) and the Caddy front-door restart handler fires. Comment only, no behaviour change. ## Apply (operator) ```bash cd ansible ansible-playbook site.yml --limit bitborg-prod,bitborg-monitoring --tags base ansible-playbook site.yml --limit bitborg-monitoring --tags monitoring ``` Then `journalctl --disk-usage` and `journalctl --header | grep -i "head realtime"` after a day to confirm rotation. ## Tested - `ansible-lint`: 0 failures, 0 warnings, 57 files. - `pnpm format:check`, `pnpm mdlint`: clean. Pre-commit hooks passed on all three commits. - Check mode as above. No apply. ## Open questions - Services host Caddy active file age (see ceiling above). - Billing database: whether a request to de-identify payment records is ever honoured is an operator decision. The runbook says this procedure does not touch it. - Legal basis for keeping `checkout_consents` and `withdrawal_requests` is my reading (consumer-law evidence, accounting trail). Confirm. - Other `auth.gitborg.se` mentions remain elsewhere in the runbook (bootstrap, Kanidm and monitoring sections). Out of scope here.
supernaut lade till 6 incheckningar 2026-10-03 00:10:30 +00:00
The privacy policy states logs are kept at most 7 days. journald had only a
size cap (SystemMaxUse), so entries could stay for weeks on a quiet host.
MaxRetentionSec=7day sets the time bound. MaxFileSec=1day rotates daily,
since journald vacuums only archived files and the default rotation is
monthly.
The access log goes to stdout, so the client IPs land in the monitoring
host's journald and are not shipped to Loki. The journald time limit from
the base role bounds them to 7 days. The comment claimed they reached Loki.
The erasure SQL skipped email_change_requests, which holds the pending new
address. It is keyed by the Kanidm sub, so step 1 now reads the sub before the
person is deleted. The Kanidm URLs used auth.gitborg.se, now auth.bitborg.se.
The section now names what is kept on purpose (checkout_consents,
withdrawal_requests, the billing database) and why.
Commands, URLs and the Grafana admin group SPN still named auth, git, www,
grafana, stats and ntfy on gitborg.se. The live Kanidm domain is
auth.bitborg.se, and grafana_admin_group derives from it. Lines that describe
the old name on purpose (the switchover history, redirects, measurements)
are unchanged.
The Grafana admin group SPN and the core-domain list in the monitoring
troubleshooting still read gitborg.se.
docs(runbook): restore the switchover note that describes the old name
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m11s
b2f498f2b2
That sentence records the group SPN before the domain rename, so it must
keep the old name.
supernaut schemalade den här ändringsförfrågan för automatisk sammanfogning när alla kontroller lyckas 2026-10-03 00:11:30 +00:00
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m11s
Obligatorisk
Detaljer
den här ändringsförfrågan är blockerad eftersom den är föråldrad.
Den här grenen är föråldrad gentemot basgrenen
Du är inte behörig att sammanfoga den här ändringsförfrågan.
Visa kommandoradsinstruktioner

Checka ut

Checka ut en ny gren från din projektkatalog och testa ändringarna.
git fetch -u origin fix/log-retention-and-erasure:fix/log-retention-and-erasure
git switch fix/log-retention-and-erasure
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!532
Ingen beskrivning angiven.