Add S3 backups to all CNPG Postgres clusters
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful

None of the CNPG clusters had any backup configured, so a lost PVC or a bad
migration meant permanent data loss. This adds continuous WAL archiving plus a
nightly base backup to Ceph RGW for every cluster, with restore docs.

- Add spec.backup.barmanObjectStore (in-tree; operator is CNPG 1.28, Barman
  Cloud Plugin not deployed) to each cnpg_cluster.yaml: WAL zstd, base bzip2,
  30-day retention, endpointCA via the reflected vault-ca-cert.
- Add cnpg_backup.yaml per app: a cephrgw ObjectStoreUser + Bucket (one
  dedicated s3://cnpg-<app> bucket and owner user per cluster, since cephrgw CRs
  and the CNPG credential Secret are namespace-scoped) and a staggered nightly
  ScheduledBackup. Credentials are minted by the operator; nothing is hardcoded.
- Add the three ceph.unkin.net CRD schemas so kubeconform can validate the CRs.
- Add docs/ (README index, cnpg-backups.md, cnpg-restore.md) covering config and
  full/PITR restore procedures.

Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
This commit is contained in:
2026-07-27 22:12:05 +10:00
parent b725bf7dcf
commit 0edbf0767f
30 changed files with 1464 additions and 0 deletions
+8
View File
@@ -0,0 +1,8 @@
# argocd-apps docs
Operational notes for the manifests in this repo.
| Doc | What it covers |
| --- | --- |
| [cnpg-backups.md](cnpg-backups.md) | How CNPG Postgres backups (WAL archiving + nightly base backups) to Ceph RGW are configured. |
| [cnpg-restore.md](cnpg-restore.md) | Restoring a CNPG cluster: full recovery, point-in-time recovery, cutover, and gotchas. |
+146
View File
@@ -0,0 +1,146 @@
# CNPG backups
Every CNPG Postgres cluster in this repo backs up to a dedicated Ceph RGW (S3)
bucket: continuous WAL archiving plus a staggered nightly base backup, 30-day
retention, compressed. Buckets and their S3 keys are provisioned by the in-estate
`cephrgw-operator` — nothing is seeded by hand. Because the operator's CRs are
namespace-scoped and CNPG reads its credential Secret from its own namespace, the
topology is **one bucket per cluster** (`s3://cnpg-<app>`), not one shared bucket.
The backup mechanism is CNPG's in-tree `barmanObjectStore` (the deployed operator
is v1.28; the Barman Cloud Plugin is not deployed — see the migration note at the
bottom).
## Where it lives
Per cluster, two files under `apps/base/<app>/`:
- `cnpg_cluster.yaml` — the `spec.backup` stanza (WAL archiving + retention).
- `cnpg_backup.yaml` — the `ObjectStoreUser` + `Bucket` (RGW provisioning) and the
nightly `ScheduledBackup`.
## The backup stanza (`spec.backup` in the Cluster)
```yaml
spec:
backup:
retentionPolicy: 30d # default; adjust per cluster
barmanObjectStore:
destinationPath: s3://cnpg-<app>
endpointURL: https://s3.ceph.unkin.net
endpointCA: # trust the internal Vault PKI CA
name: vault-ca-cert # reflected into every namespace
key: ca.crt
s3Credentials:
accessKeyId:
name: cnpg-<app>-backup-s3 # minted by the ObjectStoreUser
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-<app>-backup-s3
key: AWS_SECRET_ACCESS_KEY
serverName: <app> # path prefix inside the bucket
data:
compression: bzip2 # base backup (zstd not supported here)
jobs: 2
wal:
compression: zstd # WAL segments
maxParallel: 2
```
Setting `spec.backup.barmanObjectStore` turns on **continuous WAL archiving**
immediately (CNPG points `archive_command` at the object store). The nightly
base backup is a separate object:
## The nightly base backup (`ScheduledBackup`)
```yaml
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-<app>-nightly
namespace: <app>
spec:
schedule: "0 0 1 * * *" # 6-field cron, SECONDS first (01:00:00 daily)
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: <cluster-name>
```
Schedules are staggered so the base backups don't hit RGW at once:
| App | Cluster | Bucket | Nightly (local) |
| --- | --- | --- | --- |
| authentik | postgres | cnpg-authentik | 01:00 |
| litellm | litellm-postgres | cnpg-litellm | 01:20 |
| artifactapi | postgres | cnpg-artifactapi | 01:40 |
| woodpecker | woodpecker-postgres | cnpg-woodpecker | 02:00 |
| puppet | puppet-postgres | cnpg-puppet | 02:20 |
| encapi | postgres | cnpg-encapi | 02:40 |
| paperclip | paperclip-postgres | cnpg-paperclip | 03:00 |
| grafana | postgres | cnpg-grafana | 03:20 |
## Where the credentials come from
The `ObjectStoreUser` in `cnpg_backup.yaml` tells `cephrgw-operator` to mint an RGW
user and write its keys into a Secret; the `Bucket` makes that user the bucket owner
(full read/write on its own bucket). No keys are ever committed.
```yaml
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-<app>-backup
namespace: <app>
spec:
uid: cnpg-<app>-backup # RGW users are global; keep it unique
secretName: cnpg-<app>-backup-s3 # -> AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-<app>
namespace: <app>
spec:
bucketName: cnpg-<app>
ownerRef: cnpg-<app>-backup
retainOnDelete: true
```
Confirm the operator provisioned everything:
```bash
kubectl -n <app> get objectstoreuser,bucket
kubectl -n <app> get secret cnpg-<app>-backup-s3 \
-o jsonpath='{.data.AWS_ACCESS_KEY_ID}' | base64 -d; echo
```
## Checking backup status
```bash
# Cluster health + first-recoverability point (WAL archiving working?)
kubectl -n <app> get cluster <cluster-name> \
-o jsonpath='{.status.firstRecoverabilityPoint}{"\n"}'
# Backups taken so far
kubectl -n <app> get backups.postgresql.cnpg.io
# With the cnpg kubectl plugin (richer view, shows archiving + last backup)
kubectl cnpg status <cluster-name> -n <app>
# Kick a one-off backup right now (verify the whole path end to end)
kubectl cnpg backup <cluster-name> -n <app>
# Look at what actually landed in the bucket
aws --endpoint-url https://s3.ceph.unkin.net s3 ls s3://cnpg-<app>/<app>/
```
## Follow-up: Barman Cloud Plugin migration
CNPG 1.26+ deprecates the in-tree `barmanObjectStore` in favour of the Barman Cloud
Plugin (removal is planned for a future major). The plugin is **not** deployed today,
so this repo stays on the in-tree mechanism, which is fully functional on 1.28.
Migrating means deploying the plugin and moving each cluster's config to an
`ObjectStore` CR + `plugins:` reference — track that as separate work.
+131
View File
@@ -0,0 +1,131 @@
# CNPG restore
Recovery is always into a **new** Cluster that bootstraps from the object store —
CNPG never restores in place. The source backups live in `s3://cnpg-<app>` under
`serverName: <app>` (see [cnpg-backups.md](cnpg-backups.md)). Do all of this in the
source cluster's namespace so the `cnpg-<app>-backup-s3` Secret and `vault-ca-cert`
are present.
## (a) Full restore into a new cluster
Recover the latest available state into a fresh cluster named `postgres-restore`.
The `externalClusters` entry points at the **existing** backup path; `serverName`
under `barmanObjectStore` (the new cluster's own archive target) MUST differ from the
source, or the restored cluster will overwrite the archive it just recovered from.
```yaml
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres-restore
namespace: <app>
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:18.1-system-trixie # match source
storage:
size: 20Gi
storageClass: cephrbd-fast-delete
bootstrap:
recovery:
source: source-archive # references externalClusters below
# New archive target — DIFFERENT serverName from the source (avoids collision).
backup:
retentionPolicy: 30d
barmanObjectStore:
destinationPath: s3://cnpg-<app>
endpointURL: https://s3.ceph.unkin.net
endpointCA: {name: vault-ca-cert, key: ca.crt}
s3Credentials:
accessKeyId: {name: cnpg-<app>-backup-s3, key: AWS_ACCESS_KEY_ID}
secretAccessKey: {name: cnpg-<app>-backup-s3, key: AWS_SECRET_ACCESS_KEY}
serverName: <app>-restored # NOT "<app>"
wal: {compression: zstd}
data: {compression: bzip2}
externalClusters:
- name: source-archive
barmanObjectStore:
destinationPath: s3://cnpg-<app>
endpointURL: https://s3.ceph.unkin.net
endpointCA: {name: vault-ca-cert, key: ca.crt}
s3Credentials:
accessKeyId: {name: cnpg-<app>-backup-s3, key: AWS_ACCESS_KEY_ID}
secretAccessKey: {name: cnpg-<app>-backup-s3, key: AWS_SECRET_ACCESS_KEY}
serverName: <app> # the SOURCE archive to read from
```
```bash
kubectl apply -f postgres-restore.yaml
kubectl -n <app> get cluster postgres-restore -w # wait for Cluster in healthy state
```
## (b) Point-in-time recovery (PITR)
Same as above, but add `recoveryTarget` to stop replay at a timestamp. WAL is
replayed from the most recent base backup up to `targetTime`.
```yaml
bootstrap:
recovery:
source: source-archive
recoveryTarget:
# RFC3339 with timezone. Also valid: targetLSN, targetXID, targetName.
targetTime: "2026-07-26 14:30:00.000000+00"
```
```bash
# List backups to pick a base that precedes your target time
kubectl -n <app> get backups.postgresql.cnpg.io \
-o custom-columns=NAME:.metadata.name,START:.status.startedAt,STOP:.status.stoppedAt
```
To recover from one **specific** base backup instead of the newest, point the
source at a `Backup` object:
```yaml
externalClusters:
- name: source-archive
# ...barmanObjectStore as above...
bootstrap:
recovery:
backup:
name: <backup-object-name>
recoveryTarget:
targetTime: "2026-07-26 14:30:00+00"
```
## (c) Verify, then cut over
```bash
# 1. Sanity-check the recovered data before touching production.
kubectl cnpg psql postgres-restore -n <app> -- -c '\l'
kubectl cnpg psql postgres-restore -n <app> -d <db> -- \
-c 'select max(id), count(*) from <sanity_table>;'
# 2. Confirm the restored cluster is archiving to its NEW serverName.
kubectl cnpg status postgres-restore -n <app>
```
Cutover = repoint the app at the new cluster. CNPG service names track the Cluster
name (`<cluster>-rw` / `-ro` / `-r`), so update whatever the app connects through —
the CNPG `Pooler` (`cnpg_pooler.yaml`) `cluster.name`, or the app's DB host env — to
`postgres-restore`, then retire the old cluster. There is no in-place rename; the new
name is the cluster's identity. If you truly need the old name back, restore again
with `metadata.name` set to the original (after deleting the old one).
## (d) Gotchas
- **serverName collision.** The new cluster's `spec.backup...serverName` must differ
from the source's, or it re-uses the same path and corrupts/overwrites the source
archive on its first WAL push. Use `<app>-restored` (or similar) as above.
- **Secrets must exist in the target namespace.** `cnpg-<app>-backup-s3` and
`vault-ca-cert` are referenced by both `externalClusters` and `backup`. Restoring
into a *different* namespace means recreating (or reflecting) those first — the
`ObjectStoreUser`/`Bucket` CRs are namespace-scoped.
- **Match the image major.** Bootstrap-recovery replays WAL; use the same
`imageName` Postgres major as the source (mismatched majors will refuse to start).
- **PITR base must precede the target.** `targetTime` has to fall after a completed
base backup's start; otherwise there's nothing to replay onto.
- **`recoveryTarget` is one-shot.** It only applies during bootstrap. Once promoted,
the cluster is a normal primary — you can't "re-PITR" it; start a new restore.