Add S3 backups to all CNPG Postgres clusters
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful

None of the CNPG clusters had any backup configured, so a lost PVC or a bad
migration meant permanent data loss. This adds continuous WAL archiving plus a
nightly base backup to Ceph RGW for every cluster, with restore docs.

- Add spec.backup.barmanObjectStore (in-tree; operator is CNPG 1.28, Barman
  Cloud Plugin not deployed) to each cnpg_cluster.yaml: WAL zstd, base bzip2,
  30-day retention, endpointCA via the reflected vault-ca-cert.
- Add cnpg_backup.yaml per app: a cephrgw ObjectStoreUser + Bucket (one
  dedicated s3://cnpg-<app> bucket and owner user per cluster, since cephrgw CRs
  and the CNPG credential Secret are namespace-scoped) and a staggered nightly
  ScheduledBackup. Credentials are minted by the operator; nothing is hardcoded.
- Add the three ceph.unkin.net CRD schemas so kubeconform can validate the CRs.
- Add docs/ (README index, cnpg-backups.md, cnpg-restore.md) covering config and
  full/PITR restore procedures.

Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
This commit is contained in:
2026-07-27 22:12:05 +10:00
parent b725bf7dcf
commit 0edbf0767f
30 changed files with 1464 additions and 0 deletions
+8
View File
@@ -0,0 +1,8 @@
# argocd-apps docs
Operational notes for the manifests in this repo.
| Doc | What it covers |
| --- | --- |
| [cnpg-backups.md](cnpg-backups.md) | How CNPG Postgres backups (WAL archiving + nightly base backups) to Ceph RGW are configured. |
| [cnpg-restore.md](cnpg-restore.md) | Restoring a CNPG cluster: full recovery, point-in-time recovery, cutover, and gotchas. |