3c2bdf307a
## Why None of the 8 CNPG Postgres clusters in this repo had **any** backup configured. A lost PVC, a fat-fingered migration, or a bad app deploy meant permanent, unrecoverable data loss for authentik, litellm, artifactapi, woodpecker, puppet, encapi, paperclip and grafana. This adds continuous WAL archiving + a nightly base backup to Ceph RGW for every cluster, plus code-forward restore docs. ## What - **`spec.backup.barmanObjectStore`** on each `cnpg_cluster.yaml` — turns on continuous WAL archiving to `s3://cnpg-<app>`, WAL compressed with zstd, base backups with bzip2, 30-day retention. TLS to `s3.ceph.unkin.net` is trusted via the reflected `vault-ca-cert` (`endpointCA`). - **`cnpg_backup.yaml`** per app — a cephrgw `ObjectStoreUser` + `Bucket` (operator provisions the bucket and mints the S3 key into `cnpg-<app>-backup-s3`; **nothing is hardcoded**) and a staggered nightly `ScheduledBackup`. - **`schemas/ceph.unkin.net/*.json`** — the three cephrgw CRD schemas so kubeconform can validate the new CRs. - **`docs/`** — new docs folder (README index + `cnpg-backups.md` + `cnpg-restore.md`). ## Design decisions (answers to the open questions) **One bucket for all, or per-database?** → **Per-database (one bucket + owner user per cluster).** The cephrgw CRDs are namespace-scoped (`BucketRef`/`OwnerRef` resolve only *within the same namespace*), and CNPG reads its S3 credential Secret from its *own* namespace. A single shared bucket would require either cross-namespace bucket refs (unsupported) or hand-copying the S3 secret into all 8 namespaces (defeats "operator mints the keys"). Per-namespace `s3://cnpg-<app>` with a dedicated owner user is the simplest correct topology and needs zero manual seeding. Each user owns exactly one bucket, so owner-level (full) access is already tightly scoped — no extra `BucketAccess` grant needed. **Backup mechanism.** The deployed CNPG operator is **v1.28** (helm chart `cloudnative-pg-0.27.0`, appVersion 1.28.0). 1.26+ deprecates the in-tree `barmanObjectStore` in favour of the Barman Cloud Plugin, but the plugin is **not deployed**, and `barmanObjectStore` is still fully functional on 1.28. So this uses the in-tree mechanism. Migrating to the plugin is a follow-up (noted in `docs/cnpg-backups.md`). ## Schedule / retention (defaults — Ben to adjust) | App | Cluster | Bucket | Nightly base backup | | --- | --- | --- | --- | | authentik | postgres | cnpg-authentik | 01:00 | | litellm | litellm-postgres | cnpg-litellm | 01:20 | | artifactapi | postgres | cnpg-artifactapi | 01:40 | | woodpecker | woodpecker-postgres | cnpg-woodpecker | 02:00 | | puppet | puppet-postgres | cnpg-puppet | 02:20 | | encapi | postgres | cnpg-encapi | 02:40 | | paperclip | paperclip-postgres | cnpg-paperclip | 03:00 | | grafana | postgres | cnpg-grafana | 03:20 | Retention is **30d** across the board — flagged as a default to tune per cluster. Schedules are staggered 20 min apart so 8 base backups don't hit RGW at once. ## Validation - `kustomize build --enable-helm` + `kubeconform` (repo CI args, incl. the new ceph schemas) pass on all 8 affected overlays (paperclip validated at base — it has no overlay yet). ceph CRs resolve their schemas (`Skipped: 0`). - `pre-commit run` passes on all changed files (yamllint, no-plain-secrets, etc.). - Note: a full `ci/validate-apps.sh` run aborts locally on the unrelated `cattle-system` overlay (`chart requires kubeVersion < 1.35 vs host helm v1.36.0`) — pre-existing, reproduces on `origin/main`, unrelated to this change. ## Notes / caveats - No overlap with the woodpecker chart-bump PR (#297, overlay files only) or the logging PR (#296) beyond the three **identical** generated `schemas/ceph.unkin.net/*.json` files, which merge cleanly whichever lands first. - Credentials: no manual seeding — the cephrgw-operator mints the RGW user + keys. The only prerequisite is the operator being healthy (it is, in `cephrgw-system`). ## Follow-ups - Barman Cloud Plugin migration (deploy plugin, move clusters to `ObjectStore` CRs). - Tune per-cluster retention / schedule if the defaults don't fit. https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv Reviewed-on: #298 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net>
132 lines
5.1 KiB
Markdown
132 lines
5.1 KiB
Markdown
# CNPG restore
|
|
|
|
Recovery is always into a **new** Cluster that bootstraps from the object store —
|
|
CNPG never restores in place. The source backups live in `s3://cnpg-<app>` under
|
|
`serverName: <app>` (see [cnpg-backups.md](cnpg-backups.md)). Do all of this in the
|
|
source cluster's namespace so the `cnpg-<app>-backup-s3` Secret and `vault-ca-cert`
|
|
are present.
|
|
|
|
## (a) Full restore into a new cluster
|
|
|
|
Recover the latest available state into a fresh cluster named `postgres-restore`.
|
|
The `externalClusters` entry points at the **existing** backup path; `serverName`
|
|
under `barmanObjectStore` (the new cluster's own archive target) MUST differ from the
|
|
source, or the restored cluster will overwrite the archive it just recovered from.
|
|
|
|
```yaml
|
|
apiVersion: postgresql.cnpg.io/v1
|
|
kind: Cluster
|
|
metadata:
|
|
name: postgres-restore
|
|
namespace: <app>
|
|
spec:
|
|
instances: 3
|
|
imageName: ghcr.io/cloudnative-pg/postgresql:18.1-system-trixie # match source
|
|
storage:
|
|
size: 20Gi
|
|
storageClass: cephrbd-fast-delete
|
|
bootstrap:
|
|
recovery:
|
|
source: source-archive # references externalClusters below
|
|
|
|
# New archive target — DIFFERENT serverName from the source (avoids collision).
|
|
backup:
|
|
retentionPolicy: 30d
|
|
barmanObjectStore:
|
|
destinationPath: s3://cnpg-<app>
|
|
endpointURL: https://s3.ceph.unkin.net
|
|
endpointCA: {name: vault-ca-cert, key: ca.crt}
|
|
s3Credentials:
|
|
accessKeyId: {name: cnpg-<app>-backup-s3, key: AWS_ACCESS_KEY_ID}
|
|
secretAccessKey: {name: cnpg-<app>-backup-s3, key: AWS_SECRET_ACCESS_KEY}
|
|
serverName: <app>-restored # NOT "<app>"
|
|
wal: {compression: zstd}
|
|
data: {compression: bzip2}
|
|
|
|
externalClusters:
|
|
- name: source-archive
|
|
barmanObjectStore:
|
|
destinationPath: s3://cnpg-<app>
|
|
endpointURL: https://s3.ceph.unkin.net
|
|
endpointCA: {name: vault-ca-cert, key: ca.crt}
|
|
s3Credentials:
|
|
accessKeyId: {name: cnpg-<app>-backup-s3, key: AWS_ACCESS_KEY_ID}
|
|
secretAccessKey: {name: cnpg-<app>-backup-s3, key: AWS_SECRET_ACCESS_KEY}
|
|
serverName: <app> # the SOURCE archive to read from
|
|
```
|
|
|
|
```bash
|
|
kubectl apply -f postgres-restore.yaml
|
|
kubectl -n <app> get cluster postgres-restore -w # wait for Cluster in healthy state
|
|
```
|
|
|
|
## (b) Point-in-time recovery (PITR)
|
|
|
|
Same as above, but add `recoveryTarget` to stop replay at a timestamp. WAL is
|
|
replayed from the most recent base backup up to `targetTime`.
|
|
|
|
```yaml
|
|
bootstrap:
|
|
recovery:
|
|
source: source-archive
|
|
recoveryTarget:
|
|
# RFC3339 with timezone. Also valid: targetLSN, targetXID, targetName.
|
|
targetTime: "2026-07-26 14:30:00.000000+00"
|
|
```
|
|
|
|
```bash
|
|
# List backups to pick a base that precedes your target time
|
|
kubectl -n <app> get backups.postgresql.cnpg.io \
|
|
-o custom-columns=NAME:.metadata.name,START:.status.startedAt,STOP:.status.stoppedAt
|
|
```
|
|
|
|
To recover from one **specific** base backup instead of the newest, point the
|
|
source at a `Backup` object:
|
|
|
|
```yaml
|
|
externalClusters:
|
|
- name: source-archive
|
|
# ...barmanObjectStore as above...
|
|
bootstrap:
|
|
recovery:
|
|
backup:
|
|
name: <backup-object-name>
|
|
recoveryTarget:
|
|
targetTime: "2026-07-26 14:30:00+00"
|
|
```
|
|
|
|
## (c) Verify, then cut over
|
|
|
|
```bash
|
|
# 1. Sanity-check the recovered data before touching production.
|
|
kubectl cnpg psql postgres-restore -n <app> -- -c '\l'
|
|
kubectl cnpg psql postgres-restore -n <app> -d <db> -- \
|
|
-c 'select max(id), count(*) from <sanity_table>;'
|
|
|
|
# 2. Confirm the restored cluster is archiving to its NEW serverName.
|
|
kubectl cnpg status postgres-restore -n <app>
|
|
```
|
|
|
|
Cutover = repoint the app at the new cluster. CNPG service names track the Cluster
|
|
name (`<cluster>-rw` / `-ro` / `-r`), so update whatever the app connects through —
|
|
the CNPG `Pooler` (`cnpg_pooler.yaml`) `cluster.name`, or the app's DB host env — to
|
|
`postgres-restore`, then retire the old cluster. There is no in-place rename; the new
|
|
name is the cluster's identity. If you truly need the old name back, restore again
|
|
with `metadata.name` set to the original (after deleting the old one).
|
|
|
|
## (d) Gotchas
|
|
|
|
- **serverName collision.** The new cluster's `spec.backup...serverName` must differ
|
|
from the source's, or it re-uses the same path and corrupts/overwrites the source
|
|
archive on its first WAL push. Use `<app>-restored` (or similar) as above.
|
|
- **Secrets must exist in the target namespace.** `cnpg-<app>-backup-s3` and
|
|
`vault-ca-cert` are referenced by both `externalClusters` and `backup`. Restoring
|
|
into a *different* namespace means recreating (or reflecting) those first — the
|
|
`ObjectStoreUser`/`Bucket` CRs are namespace-scoped.
|
|
- **Match the image major.** Bootstrap-recovery replays WAL; use the same
|
|
`imageName` Postgres major as the source (mismatched majors will refuse to start).
|
|
- **PITR base must precede the target.** `targetTime` has to fall after a completed
|
|
base backup's start; otherwise there's nothing to replay onto.
|
|
- **`recoveryTarget` is one-shot.** It only applies during bootstrap. Once promoted,
|
|
the cluster is a normal primary — you can't "re-PITR" it; start a new restore.
|