Add S3 backups to all CNPG Postgres clusters #298

Merged
benvin merged 1 commits from benvin/cnpg-backups into main 2026-07-27 23:57:09 +10:00
Owner

Why

None of the 8 CNPG Postgres clusters in this repo had any backup configured. A
lost PVC, a fat-fingered migration, or a bad app deploy meant permanent, unrecoverable
data loss for authentik, litellm, artifactapi, woodpecker, puppet, encapi, paperclip
and grafana. This adds continuous WAL archiving + a nightly base backup to Ceph RGW for
every cluster, plus code-forward restore docs.

What

  • spec.backup.barmanObjectStore on each cnpg_cluster.yaml — turns on continuous
    WAL archiving to s3://cnpg-<app>, WAL compressed with zstd, base backups with bzip2,
    30-day retention. TLS to s3.ceph.unkin.net is trusted via the reflected
    vault-ca-cert (endpointCA).
  • cnpg_backup.yaml per app — a cephrgw ObjectStoreUser + Bucket (operator
    provisions the bucket and mints the S3 key into cnpg-<app>-backup-s3; nothing is
    hardcoded
    ) and a staggered nightly ScheduledBackup.
  • schemas/ceph.unkin.net/*.json — the three cephrgw CRD schemas so kubeconform can
    validate the new CRs.
  • docs/ — new docs folder (README index + cnpg-backups.md + cnpg-restore.md).

Design decisions (answers to the open questions)

One bucket for all, or per-database?Per-database (one bucket + owner user per
cluster).
The cephrgw CRDs are namespace-scoped (BucketRef/OwnerRef resolve only
within the same namespace), and CNPG reads its S3 credential Secret from its own
namespace. A single shared bucket would require either cross-namespace bucket refs
(unsupported) or hand-copying the S3 secret into all 8 namespaces (defeats "operator
mints the keys"). Per-namespace s3://cnpg-<app> with a dedicated owner user is the
simplest correct topology and needs zero manual seeding. Each user owns exactly one
bucket, so owner-level (full) access is already tightly scoped — no extra BucketAccess
grant needed.

Backup mechanism. The deployed CNPG operator is v1.28 (helm chart
cloudnative-pg-0.27.0, appVersion 1.28.0). 1.26+ deprecates the in-tree
barmanObjectStore in favour of the Barman Cloud Plugin, but the plugin is not
deployed
, and barmanObjectStore is still fully functional on 1.28. So this uses the
in-tree mechanism. Migrating to the plugin is a follow-up (noted in docs/cnpg-backups.md).

Schedule / retention (defaults — Ben to adjust)

App Cluster Bucket Nightly base backup
authentik postgres cnpg-authentik 01:00
litellm litellm-postgres cnpg-litellm 01:20
artifactapi postgres cnpg-artifactapi 01:40
woodpecker woodpecker-postgres cnpg-woodpecker 02:00
puppet puppet-postgres cnpg-puppet 02:20
encapi postgres cnpg-encapi 02:40
paperclip paperclip-postgres cnpg-paperclip 03:00
grafana postgres cnpg-grafana 03:20

Retention is 30d across the board — flagged as a default to tune per cluster.
Schedules are staggered 20 min apart so 8 base backups don't hit RGW at once.

Validation

  • kustomize build --enable-helm + kubeconform (repo CI args, incl. the new ceph
    schemas) pass on all 8 affected overlays (paperclip validated at base — it has no
    overlay yet). ceph CRs resolve their schemas (Skipped: 0).
  • pre-commit run passes on all changed files (yamllint, no-plain-secrets, etc.).
  • Note: a full ci/validate-apps.sh run aborts locally on the unrelated
    cattle-system overlay (chart requires kubeVersion < 1.35 vs host helm v1.36.0) —
    pre-existing, reproduces on origin/main, unrelated to this change.

Notes / caveats

  • No overlap with the woodpecker chart-bump PR (#297, overlay files only) or the logging
    PR (#296) beyond the three identical generated schemas/ceph.unkin.net/*.json
    files, which merge cleanly whichever lands first.
  • Credentials: no manual seeding — the cephrgw-operator mints the RGW user + keys. The
    only prerequisite is the operator being healthy (it is, in cephrgw-system).

Follow-ups

  • Barman Cloud Plugin migration (deploy plugin, move clusters to ObjectStore CRs).
  • Tune per-cluster retention / schedule if the defaults don't fit.

https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv

## Why None of the 8 CNPG Postgres clusters in this repo had **any** backup configured. A lost PVC, a fat-fingered migration, or a bad app deploy meant permanent, unrecoverable data loss for authentik, litellm, artifactapi, woodpecker, puppet, encapi, paperclip and grafana. This adds continuous WAL archiving + a nightly base backup to Ceph RGW for every cluster, plus code-forward restore docs. ## What - **`spec.backup.barmanObjectStore`** on each `cnpg_cluster.yaml` — turns on continuous WAL archiving to `s3://cnpg-<app>`, WAL compressed with zstd, base backups with bzip2, 30-day retention. TLS to `s3.ceph.unkin.net` is trusted via the reflected `vault-ca-cert` (`endpointCA`). - **`cnpg_backup.yaml`** per app — a cephrgw `ObjectStoreUser` + `Bucket` (operator provisions the bucket and mints the S3 key into `cnpg-<app>-backup-s3`; **nothing is hardcoded**) and a staggered nightly `ScheduledBackup`. - **`schemas/ceph.unkin.net/*.json`** — the three cephrgw CRD schemas so kubeconform can validate the new CRs. - **`docs/`** — new docs folder (README index + `cnpg-backups.md` + `cnpg-restore.md`). ## Design decisions (answers to the open questions) **One bucket for all, or per-database?** → **Per-database (one bucket + owner user per cluster).** The cephrgw CRDs are namespace-scoped (`BucketRef`/`OwnerRef` resolve only *within the same namespace*), and CNPG reads its S3 credential Secret from its *own* namespace. A single shared bucket would require either cross-namespace bucket refs (unsupported) or hand-copying the S3 secret into all 8 namespaces (defeats "operator mints the keys"). Per-namespace `s3://cnpg-<app>` with a dedicated owner user is the simplest correct topology and needs zero manual seeding. Each user owns exactly one bucket, so owner-level (full) access is already tightly scoped — no extra `BucketAccess` grant needed. **Backup mechanism.** The deployed CNPG operator is **v1.28** (helm chart `cloudnative-pg-0.27.0`, appVersion 1.28.0). 1.26+ deprecates the in-tree `barmanObjectStore` in favour of the Barman Cloud Plugin, but the plugin is **not deployed**, and `barmanObjectStore` is still fully functional on 1.28. So this uses the in-tree mechanism. Migrating to the plugin is a follow-up (noted in `docs/cnpg-backups.md`). ## Schedule / retention (defaults — Ben to adjust) | App | Cluster | Bucket | Nightly base backup | | --- | --- | --- | --- | | authentik | postgres | cnpg-authentik | 01:00 | | litellm | litellm-postgres | cnpg-litellm | 01:20 | | artifactapi | postgres | cnpg-artifactapi | 01:40 | | woodpecker | woodpecker-postgres | cnpg-woodpecker | 02:00 | | puppet | puppet-postgres | cnpg-puppet | 02:20 | | encapi | postgres | cnpg-encapi | 02:40 | | paperclip | paperclip-postgres | cnpg-paperclip | 03:00 | | grafana | postgres | cnpg-grafana | 03:20 | Retention is **30d** across the board — flagged as a default to tune per cluster. Schedules are staggered 20 min apart so 8 base backups don't hit RGW at once. ## Validation - `kustomize build --enable-helm` + `kubeconform` (repo CI args, incl. the new ceph schemas) pass on all 8 affected overlays (paperclip validated at base — it has no overlay yet). ceph CRs resolve their schemas (`Skipped: 0`). - `pre-commit run` passes on all changed files (yamllint, no-plain-secrets, etc.). - Note: a full `ci/validate-apps.sh` run aborts locally on the unrelated `cattle-system` overlay (`chart requires kubeVersion < 1.35 vs host helm v1.36.0`) — pre-existing, reproduces on `origin/main`, unrelated to this change. ## Notes / caveats - No overlap with the woodpecker chart-bump PR (#297, overlay files only) or the logging PR (#296) beyond the three **identical** generated `schemas/ceph.unkin.net/*.json` files, which merge cleanly whichever lands first. - Credentials: no manual seeding — the cephrgw-operator mints the RGW user + keys. The only prerequisite is the operator being healthy (it is, in `cephrgw-system`). ## Follow-ups - Barman Cloud Plugin migration (deploy plugin, move clusters to `ObjectStore` CRs). - Tune per-cluster retention / schedule if the defaults don't fit. https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
unkinben added 1 commit 2026-07-27 22:13:14 +10:00
Add S3 backups to all CNPG Postgres clusters
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
0edbf0767f
None of the CNPG clusters had any backup configured, so a lost PVC or a bad
migration meant permanent data loss. This adds continuous WAL archiving plus a
nightly base backup to Ceph RGW for every cluster, with restore docs.

- Add spec.backup.barmanObjectStore (in-tree; operator is CNPG 1.28, Barman
  Cloud Plugin not deployed) to each cnpg_cluster.yaml: WAL zstd, base bzip2,
  30-day retention, endpointCA via the reflected vault-ca-cert.
- Add cnpg_backup.yaml per app: a cephrgw ObjectStoreUser + Bucket (one
  dedicated s3://cnpg-<app> bucket and owner user per cluster, since cephrgw CRs
  and the CNPG credential Secret are namespace-scoped) and a staggered nightly
  ScheduledBackup. Credentials are minted by the operator; nothing is hardcoded.
- Add the three ceph.unkin.net CRD schemas so kubeconform can validate the CRs.
- Add docs/ (README index, cnpg-backups.md, cnpg-restore.md) covering config and
  full/PITR restore procedures.

Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
benvin merged commit 3c2bdf307a into main 2026-07-27 23:57:09 +10:00
benvin deleted branch benvin/cnpg-backups 2026-07-27 23:57:09 +10:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: unkin/argocd-apps#298