HA: PostgreSQL replication / managed HA pair #43

Öppen
öppnade 2026-07-09 23:03:57 +00:00 av supernaut · 1 kommentar
Ägare

Move PostgreSQL to a managed/replicated setup or a dedicated HA pair.

Epic: gitborg/gitborg-docs#4

Move PostgreSQL to a managed/replicated setup or a dedicated HA pair. Epic: gitborg/gitborg-docs#4
Upphovsperson
Ägare

Design/options doc: tracked internally.

TL;DR — defer active HA; harden the single instance now. We run one containerised Postgres 17 (Forgejo + web) on one Bahnhof VM. Recovery today: daily age-encrypted dumps (ADR 0022/0011) proven by the weekly restore drill (ADR 0027) → ~RPO 24h, RTO hours. HA buys lower RTO (mins) and, with streaming replication, lower RPO (secs) — but the DB is only one SPOF on a single-VM stack, scale is tiny, and there's no SLA yet. Backups suffice for now.

Options weighed:

  • (a) Hardened single [recommended now] — add WAL archiving (RPO→mins), a drilled promotion/restore runbook. ~0 cost, stays fully sovereign.
  • (b) Streaming standby VM (repmgr, later Patroni) — real HA, but ~2× DB infra + failover-correctness burden; revisit post-onboarding, inside the whole-stack HA epic.
  • (c) Managed sovereign — Bahnhof has no DBaaS. Glesys (our SE backup provider) offers managed PG18 but no replication/failover/PITR → managed-but-not-HA. Adminor (Stockholm) is auto-failover HA but needs a cross-provider VLAN/VPN tunnel from Bahnhof + adds a processor. No clean sovereign managed-HA fit today. Never AWS/GCP/Azure (principle 1).
  • (d) CloudNativePG/Stolon — needs Kubernetes (contra ADR 0002); overkill. Rejected.

Path: WAL archiving + drilled runbook now → streaming replication when users/SLA warrant → managed only if a sovereign co-located HA option appears. Suggest also making the Forgejo/web DB host a variable (currently hard-coded postgres:5432). Recommend moving off needs-triage.

**Design/options doc:** tracked internally. **TL;DR — defer active HA; harden the single instance now.** We run one containerised Postgres 17 (Forgejo + web) on one Bahnhof VM. Recovery today: daily age-encrypted dumps (ADR 0022/0011) proven by the weekly restore drill (ADR 0027) → ~RPO 24h, RTO hours. HA buys lower RTO (mins) and, with streaming replication, lower RPO (secs) — but the DB is only **one** SPOF on a single-VM stack, scale is tiny, and there's no SLA yet. Backups suffice for now. **Options weighed:** - **(a) Hardened single [recommended now]** — add WAL archiving (RPO→mins), a drilled promotion/restore runbook. ~0 cost, stays fully sovereign. - **(b) Streaming standby VM (repmgr, later Patroni)** — real HA, but ~2× DB infra + failover-correctness burden; revisit post-onboarding, inside the whole-stack HA epic. - **(c) Managed sovereign** — Bahnhof has no DBaaS. **Glesys** (our SE backup provider) offers managed PG18 but **no replication/failover/PITR** → managed-but-not-HA. **Adminor** (Stockholm) *is* auto-failover HA but needs a **cross-provider VLAN/VPN tunnel** from Bahnhof + adds a processor. No clean sovereign managed-HA fit today. Never AWS/GCP/Azure (principle 1). - **(d) CloudNativePG/Stolon** — needs Kubernetes (contra ADR 0002); overkill. Rejected. **Path:** WAL archiving + drilled runbook now → streaming replication when users/SLA warrant → managed only if a sovereign co-located HA option appears. Suggest also making the Forgejo/web DB host a variable (currently hard-coded `postgres:5432`). Recommend moving off `needs-triage`.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#43
Ingen beskrivning angiven.