None of the CNPG clusters had any backup configured, so a lost PVC or a bad migration meant permanent data loss. This adds continuous WAL archiving plus a nightly base backup to Ceph RGW for every cluster, with restore docs. - Add spec.backup.barmanObjectStore (in-tree; operator is CNPG 1.28, Barman Cloud Plugin not deployed) to each cnpg_cluster.yaml: WAL zstd, base bzip2, 30-day retention, endpointCA via the reflected vault-ca-cert. - Add cnpg_backup.yaml per app: a cephrgw ObjectStoreUser + Bucket (one dedicated s3://cnpg-<app> bucket and owner user per cluster, since cephrgw CRs and the CNPG credential Secret are namespace-scoped) and a staggered nightly ScheduledBackup. Credentials are minted by the operator; nothing is hardcoded. - Add the three ceph.unkin.net CRD schemas so kubeconform can validate the CRs. - Add docs/ (README index, cnpg-backups.md, cnpg-restore.md) covering config and full/PITR restore procedures. Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
5.1 KiB
CNPG restore
Recovery is always into a new Cluster that bootstraps from the object store —
CNPG never restores in place. The source backups live in s3://cnpg-<app> under
serverName: <app> (see cnpg-backups.md). Do all of this in the
source cluster's namespace so the cnpg-<app>-backup-s3 Secret and vault-ca-cert
are present.
(a) Full restore into a new cluster
Recover the latest available state into a fresh cluster named postgres-restore.
The externalClusters entry points at the existing backup path; serverName
under barmanObjectStore (the new cluster's own archive target) MUST differ from the
source, or the restored cluster will overwrite the archive it just recovered from.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres-restore
namespace: <app>
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:18.1-system-trixie # match source
storage:
size: 20Gi
storageClass: cephrbd-fast-delete
bootstrap:
recovery:
source: source-archive # references externalClusters below
# New archive target — DIFFERENT serverName from the source (avoids collision).
backup:
retentionPolicy: 30d
barmanObjectStore:
destinationPath: s3://cnpg-<app>
endpointURL: https://s3.ceph.unkin.net
endpointCA: {name: vault-ca-cert, key: ca.crt}
s3Credentials:
accessKeyId: {name: cnpg-<app>-backup-s3, key: AWS_ACCESS_KEY_ID}
secretAccessKey: {name: cnpg-<app>-backup-s3, key: AWS_SECRET_ACCESS_KEY}
serverName: <app>-restored # NOT "<app>"
wal: {compression: zstd}
data: {compression: bzip2}
externalClusters:
- name: source-archive
barmanObjectStore:
destinationPath: s3://cnpg-<app>
endpointURL: https://s3.ceph.unkin.net
endpointCA: {name: vault-ca-cert, key: ca.crt}
s3Credentials:
accessKeyId: {name: cnpg-<app>-backup-s3, key: AWS_ACCESS_KEY_ID}
secretAccessKey: {name: cnpg-<app>-backup-s3, key: AWS_SECRET_ACCESS_KEY}
serverName: <app> # the SOURCE archive to read from
kubectl apply -f postgres-restore.yaml
kubectl -n <app> get cluster postgres-restore -w # wait for Cluster in healthy state
(b) Point-in-time recovery (PITR)
Same as above, but add recoveryTarget to stop replay at a timestamp. WAL is
replayed from the most recent base backup up to targetTime.
bootstrap:
recovery:
source: source-archive
recoveryTarget:
# RFC3339 with timezone. Also valid: targetLSN, targetXID, targetName.
targetTime: "2026-07-26 14:30:00.000000+00"
# List backups to pick a base that precedes your target time
kubectl -n <app> get backups.postgresql.cnpg.io \
-o custom-columns=NAME:.metadata.name,START:.status.startedAt,STOP:.status.stoppedAt
To recover from one specific base backup instead of the newest, point the
source at a Backup object:
externalClusters:
- name: source-archive
# ...barmanObjectStore as above...
bootstrap:
recovery:
backup:
name: <backup-object-name>
recoveryTarget:
targetTime: "2026-07-26 14:30:00+00"
(c) Verify, then cut over
# 1. Sanity-check the recovered data before touching production.
kubectl cnpg psql postgres-restore -n <app> -- -c '\l'
kubectl cnpg psql postgres-restore -n <app> -d <db> -- \
-c 'select max(id), count(*) from <sanity_table>;'
# 2. Confirm the restored cluster is archiving to its NEW serverName.
kubectl cnpg status postgres-restore -n <app>
Cutover = repoint the app at the new cluster. CNPG service names track the Cluster
name (<cluster>-rw / -ro / -r), so update whatever the app connects through —
the CNPG Pooler (cnpg_pooler.yaml) cluster.name, or the app's DB host env — to
postgres-restore, then retire the old cluster. There is no in-place rename; the new
name is the cluster's identity. If you truly need the old name back, restore again
with metadata.name set to the original (after deleting the old one).
(d) Gotchas
- serverName collision. The new cluster's
spec.backup...serverNamemust differ from the source's, or it re-uses the same path and corrupts/overwrites the source archive on its first WAL push. Use<app>-restored(or similar) as above. - Secrets must exist in the target namespace.
cnpg-<app>-backup-s3andvault-ca-certare referenced by bothexternalClustersandbackup. Restoring into a different namespace means recreating (or reflecting) those first — theObjectStoreUser/BucketCRs are namespace-scoped. - Match the image major. Bootstrap-recovery replays WAL; use the same
imageNamePostgres major as the source (mismatched majors will refuse to start). - PITR base must precede the target.
targetTimehas to fall after a completed base backup's start; otherwise there's nothing to replay onto. recoveryTargetis one-shot. It only applies during bootstrap. Once promoted, the cluster is a normal primary — you can't "re-PITR" it; start a new restore.