- Jinja 51.2%
- Python 28.7%
- Shell 8.9%
- HCL 8%
- CSS 1.2%
- Övrigt 2%
| Filnamn | Senaste incheckningsmeddelande | Senaste incheckningsdatum |
|---|---|---|
|
Alla kontroller lyckades
ci / ci (push) Successful in 1m56s
## What
Remove `reconciler_rename_transition_exempt` and its concatenation into `reconciler_user_exempt_extra` in `ansible/group_vars/all/vars.yml`. Fix the comment on `forgejo_service_accounts` that pointed at it.
## Why
The list was a temporary bridge for the service-account rename. Its own comment said to remove it when the tranche closed. `forgejo_service_accounts` already declares the seven `bitborg-*` accounts, and the role default `reconciler_user_exempt` derives from it.
The reconciler is authoritative over quota: an account missing from `USER_EXEMPT` is demoted to `participant` within one tick. So the removal was checked against live state, not assumed.
## Consumers traced
- `reconciler_rename_transition_exempt`: only `reconciler_user_exempt_extra` (vars.yml). No other reference in `ansible/`, `docs/`, `scripts/`.
- `reconciler_user_exempt_extra` feeds `reconciler_user_exempt` (role default) and then `USER_EXEMPT` in `reconciler.env.j2`. The reconciler reads it as a plain list (`parseList(USER_EXEMPT)`, exact-name set lookup).
- All seven `bitborg-*` names the list added are already in `forgejo_service_accounts`.
- No `gitborg-*` service account exists on the server: `fj --host git.bitborg.se user search` returns no match for all seven. The `bitborg-*` names match (negative control). Legacy names in `USER_EXEMPT` were inert.
## Check-mode proof
`ansible-playbook site.yml --check --diff --tags reconciler --limit bitborg-prod` (from `ansible/`)
| | ok | changed |
| --- | --- | --- |
| main | 27 | 0 |
| this branch | 27 | 1 |
The one change is `reconciler : Install the reconciler environment file (full contract)`. The diff itself is hidden (the file holds secrets, task is `no_log`), so the rendered list was evaluated directly with `ansible -m debug` against the same inventory, before and after:
```
before: 9 base names + 14 transition names (7 gitborg-*, 7 bitborg-* duplicates)
after: test-passkey-check, bitborg-renovate, bitborg-reconciler, bitborg-webhook-admin,
bitborg-runner-controller, bitborg-ci, bitborg-bot, bitborg-token-audit, <operator admin>
```
The "after" set equals the "before" set minus the 14 transition entries. Every live service account and the operator admin remain.
## Other checks
- `ansible-lint`: passed, 0 failures, 0 warnings (206 files).
- lefthook pre-commit (prettier): passed.
## Apply (operator)
```
cd ansible
ansible-playbook site.yml --tags reconciler --limit bitborg-prod
```
Expected: `reconciler : Install the reconciler environment file` changed, nothing else. The timer picks up the new env on its next run.
Verify after the next run (Loki, `bitborg-grafana`):
- The reconciler run log line appears and reports no demotions (no account moved to `participant`, no unexpected group removals).
- Spot check that `bitborg-ci` and `bitborg-renovate` can still push (no 413).
Rollback: revert this commit and re-apply.
Reviewed-on: #531
|
||
| .forgejo/workflows | ||
| .vscode | ||
| ansible | ||
| assets | ||
| cost | ||
| docs | ||
| local | ||
| opentofu | ||
| scripts | ||
| .editorconfig | ||
| .gitignore | ||
| .markdownlint-cli2.jsonc | ||
| .prettierignore | ||
| CONTRIBUTING.md | ||
| lefthook.json | ||
| package.json | ||
| pnpm-lock.yaml | ||
| pnpm-workspace.yaml | ||
| README.md | ||
| renovate.json | ||
| ruff.toml | ||
bitborg-infra
Infrastructure-as-code and decision record for bitborg, a self-hosted git hosting service built on Forgejo and running on Bahnhof Cloud (Swedish, GDPR-focused).
This repository is intended to eventually be hosted on bitborg itself — the service hosts its own infrastructure code.
Guiding principles
bitborg is built, in priority order, to be:
- Owned, operated and hosted in Sweden — Swedish-owned, Swedish-operated, on Swedish soil. Users' code never passes through US-owned or US-operated infrastructure (beyond reach of the US CLOUD Act).
- Respectful of user integrity — GDPR and other EU/EEA data-protection law today, and built to adapt to future regulation.
- Free and open source — a fully FOSS stack, preferring the freer alternative (copyleft, community/non-profit governance) when options are comparable.
- Secure with users' data — minimal attack surface, least privilege, secrets never in plaintext, encrypted backups and tested restores.
These take precedence over convenience and govern every decision in this repo. See
principles.md (bitborg-docs) for what each means concretely and how the
stack embodies them.
What this is
A single instance on Bahnhof's OpenStack VPC runs Forgejo + PostgreSQL + Kanidm
(identity provider) + the public portal (bitborg-web) as rootless
Podman containers managed by Quadlet (systemd), fronted by Caddy for automatic
TLS across four hosts. The instance is provisioned with OpenTofu and configured with
Ansible; a final health-check role gates every apply, probing the core services before
the run is considered done. Kanidm is the identity provider (OIDC) and source of truth for tier entitlements;
a reconciler systemd timer applies those entitlements to Forgejo (LFS quota, org-create,
and CI/Actions access). CI/CD (Forgejo Actions) is tiered and uses ephemeral
single-use OpenStack VMs — one VM per job, booted and reaped by a runner-controller role
on the services host (ADR 0021). There is no standing runner VM. bitborg-web deploys through
that pipeline — a merge to its protected main triggers an ephemeral ci runner that builds
an image, pushes it to the Forgejo registry, and the services host pulls it via podman auto-update (CI never touches prod; ADR 0019). Backups run on systemd timers as age-encrypted
archives on a dedicated Cinder volume (/srv/gitborg-backup; ADR 0022), with off-site upload
enabled to Glesys Object Storage (Swedish, Falkenberg) as primary and Hetzner Object Storage (hel1, EU) as secondary redundancy
(ADR 0011), and a weekly backup-drill role restores the latest off-site archive onto a
throwaway VM to prove the backup actually recovers. Dependency updates are automated by
self-hosted Renovate (daily timer; ADR 0023), with a registry-retention role pruning old
CI images out of the Forgejo registry and a token-audit role tracking service-account token
age. Web logins are SSO-only through Kanidm ("Sign in
with Bitborg Auth"; ADR 0014). Observability — metrics, logs, Grafana dashboards, alerting
(ntfy + email), and cookieless GoatCounter analytics — runs on a separate, opt-in
monitoring VM that the services host only pushes to, with mutual down-detection between the
two (ADR 0020).
flowchart LR
Internet([Internet])
Admin([Admin])
Internet -->|"bitborg.se :443"| Caddy
Internet -->|"www :443"| Caddy
Internet -->|"git :443"| Caddy
Internet -->|"auth :443"| Caddy
Internet -->|":22 git SSH"| Forgejo["Forgejo<br/>:3000 + SSH"]
Admin -->|":2222"| SSHD["admin sshd"]
Caddy -->|"apex → 308 www"| Caddy
Caddy -->|www| Web["bitborg-web<br/>Astro"]
Caddy -->|git| Forgejo
Caddy -->|"auth (re-encrypt)"| Kanidm["Kanidm<br/>OIDC IdP :8443"]
Web --> Forgejo
Forgejo -.->|"OIDC login"| Kanidm
Web --> Postgres[("PostgreSQL")]
Forgejo --> Postgres
Reconciler["reconciler<br/>(systemd timer)"]
Reconciler -.->|"read entitlements"| Kanidm
Reconciler -.->|"apply quota / org-create / Actions"| Forgejo
Forgejo -.-> Timers["systemd timers"]
Timers --> Local["age-encrypted backups<br/>(dedicated volume /srv/gitborg-backup)"]
Local -->|"off-site (enabled)"| S3[("Glesys S3 (primary)<br/>+ Hetzner S3 (secondary)<br/>hel1 · interim EU")]
Forgejo -.->|"OIDC (Bitborg Auth)"| Kanidm
Renovate["renovate<br/>(daily timer, ADR 0023)"]
Renovate -.->|"open update PRs"| Forgejo
RunCtl["runner-controller<br/>(ADR 0021, on services host)"]
EphVM[/"ephemeral runner VM<br/>(single-use, label: ci)"\]
RunCtl -.->|"boot / reap"| EphVM
EphVM -.->|"Actions, git :443"| Caddy
Agents["monitoring-agent<br/>(push)"] -.->|"metrics + logs + cross-probe"| Mon["Monitoring VM<br/>(opt-in, ADR 0020)<br/>VictoriaMetrics · Loki · Grafana<br/>alerts → ntfy/email · GoatCounter"]
Internet -->|"grafana/stats/ntfy :443"| Mon
Mon -.->|"OIDC (Bitborg Auth)"| Kanidm
The end-to-end flows — SSO login (Forgejo + Grafana via "Bitborg Auth"), Kanidm
provisioning, reconciliation, CI/CD deploy, backups, and
observability/alerting — are diagrammed in
architecture.md (bitborg-docs).
Layout
| Path | Purpose |
|---|---|
principles.md (bitborg-docs) |
Guiding principles that govern every decision |
architecture.md (bitborg-docs) |
Component overview and how they fit together |
docs/runbook.md |
Operations: deploy, upgrade, restore, rotate secrets |
docs/roadmap.md |
Not-yet-built growth: HA, payments |
decisions/ (bitborg-docs) |
Architecture Decision Records (ADRs) |
ansible/ |
All host provisioning and service configuration |
opentofu/ |
OpenTofu config to provision the host on OpenStack |
cost/ |
Bahnhof unit prices + gathered usage samples |
scripts/ |
Cost audit/30-day projection and complete teardown |
local/ |
Podman preview of Forgejo config/branding (no cloud) |
Git hooks
lefthook runs two hooks. pre-commit formats and fixes staged files. pre-push runs the same
static checks CI runs, so a lint failure is caught before it costs a CI round trip:
| pre-push job | CI equivalent |
|---|---|
| format (prettier) | pnpm format:check |
| lint markdown | pnpm mdlint |
| lint python (ruff) | ./scripts/ruff.sh |
| shellcheck | ./scripts/shellcheck.sh |
| renovate annotations are live | ./scripts/check-renovate-annotations.py |
| metric names are consistent | ./scripts/check-metric-names.py |
| runner-controller logic tests | test_controller_logic.py |
| opentofu fmt | tofu fmt -check -recursive |
| ansible-lint | ansible-lint (same throwaway vault password as CI) |
| secret scan (gitleaks) | gitleaks git --redact --no-banner . |
Every job carries skip_empty: false, and that is load-bearing. Lefthook decides which files a push
touches by diffing against the branch's remote counterpart. A brand-new branch has none, so the
file list is empty and every job is skipped with "no matching push files" — which is exactly the
case a feature branch is in the first time it is pushed. Measured: without it, a push carrying a
deliberate ruff E501 violation sailed straight through. With it, the same push is rejected.
Three deliberate differences, all worth knowing:
- The smoke tests are NOT in the hook.
pnpm smoke:containersandpnpm smoke:textfiledo realpodman runcalls with image pulls. On every push that is slow enough that people would start skipping the hook, which costs more than it saves. CI still gates them, so a smoke failure is caught there rather than here. Run them by hand when touching*.container.j2hardening. - The alert-rule check is NOT in the hook either.
./scripts/check-alert-rules.pyparses the rendered vmalert rules with vmalert itself, which means a container run and an image pull, so it belongs with the smoke tests rather than in a per-push hook. Run it by hand when touchingalert-rules.yml.j2or the thresholds in the monitoring role's defaults. Its--self-testproves it can still fail. - A missing tool warns, it does not fail. If
tofu,ansible-lintorgitleaksis not installed the job prints a loud warning and passes, because blocking a push on a tool you have not installed is worse than the round trip. The warning says plainly that CI will still run the check.
Bypass with LEFTHOOK=0 git push when you genuinely need to. CI is the real gate either way.
Quickstart
Prerequisites: a Bahnhof OpenStack project with
clouds.yamlconfigured, OpenTofu and Ansible installed locally, and theansible-vaultpassword for the encrypted secrets. The image is Debian 13 (trixie)+ (needs Podman ≥ 4.4 for Quadlet) — already the default in the OpenTofu config.
# 1. Provision the instance on Bahnhof OpenStack
cd opentofu
cp terraform.tfvars.example terraform.tfvars # set ssh_public_key
tofu init && tofu apply # outputs the floating IP
# 2. Configure it (copy the floating IP into the inventory first)
cd ../ansible
ansible-galaxy install -r requirements.yml # install collections
# set the floating IP in inventory/hosts.yml and your domain in group_vars/all/vars.yml
ansible-playbook site.yml # SSH port is auto-detected (see runbook §Initial deploy)
Then complete Forgejo first-run (create admin, lock down registration) per
docs/runbook.md.
To preview configuration and branding changes locally — no cloud resources or vault
needed: cd local && make up (see local/README.md).
Decisions at a glance
| Area | Choice |
|---|---|
| Hosting | Bahnhof OpenStack VPC (OpenTofu) |
| Container runtime | Podman + Quadlet (systemd) |
| Config management | Ansible |
| Database | PostgreSQL (container) |
| Reverse proxy / TLS | Caddy (automatic HTTPS, Let's Encrypt) |
| Identity / SSO | Kanidm (OIDC) on auth; group→role mapping (ADR 0014) |
| Public portal | bitborg-web (Astro) on www; open sign-up (ADR 0029) |
| CI/CD | Forgejo Actions; ephemeral single-use VMs per job (label ci); runner-controller on services host; tiered, reconciler-gated (ADR 0021) |
| Deploy (bitborg-web) | Ephemeral ci runner builds image → Forgejo registry → host pulls via auto-update (ADR 0019/0021) |
| Observability | Separate, opt-in VM: VictoriaMetrics + Loki + Grafana + alerts (ntfy/email); push-only agents; cookieless GoatCounter analytics (ADR 0020) |
| Backups | age-encrypted archives on a dedicated Cinder volume (/srv/gitborg-backup); off-site to Glesys S3 (Falkenberg, primary) and Hetzner S3 (hel1, secondary, interim EU) (ADR 0022) |
| Dependency updates | Self-hosted Renovate (daily timer); pinned images mirrored into the Forgejo registry (ADR 0023) |
| Service accounts | Scoped Forgejo service accounts for reconciler/renovate/CI/bot (ADR 0024) |
| Email (deferred) | Sweego (EU); transactional + verification |
See docs/decisions/ for the full reasoning behind each.