3c2bdf307a
## Why None of the 8 CNPG Postgres clusters in this repo had **any** backup configured. A lost PVC, a fat-fingered migration, or a bad app deploy meant permanent, unrecoverable data loss for authentik, litellm, artifactapi, woodpecker, puppet, encapi, paperclip and grafana. This adds continuous WAL archiving + a nightly base backup to Ceph RGW for every cluster, plus code-forward restore docs. ## What - **`spec.backup.barmanObjectStore`** on each `cnpg_cluster.yaml` — turns on continuous WAL archiving to `s3://cnpg-<app>`, WAL compressed with zstd, base backups with bzip2, 30-day retention. TLS to `s3.ceph.unkin.net` is trusted via the reflected `vault-ca-cert` (`endpointCA`). - **`cnpg_backup.yaml`** per app — a cephrgw `ObjectStoreUser` + `Bucket` (operator provisions the bucket and mints the S3 key into `cnpg-<app>-backup-s3`; **nothing is hardcoded**) and a staggered nightly `ScheduledBackup`. - **`schemas/ceph.unkin.net/*.json`** — the three cephrgw CRD schemas so kubeconform can validate the new CRs. - **`docs/`** — new docs folder (README index + `cnpg-backups.md` + `cnpg-restore.md`). ## Design decisions (answers to the open questions) **One bucket for all, or per-database?** → **Per-database (one bucket + owner user per cluster).** The cephrgw CRDs are namespace-scoped (`BucketRef`/`OwnerRef` resolve only *within the same namespace*), and CNPG reads its S3 credential Secret from its *own* namespace. A single shared bucket would require either cross-namespace bucket refs (unsupported) or hand-copying the S3 secret into all 8 namespaces (defeats "operator mints the keys"). Per-namespace `s3://cnpg-<app>` with a dedicated owner user is the simplest correct topology and needs zero manual seeding. Each user owns exactly one bucket, so owner-level (full) access is already tightly scoped — no extra `BucketAccess` grant needed. **Backup mechanism.** The deployed CNPG operator is **v1.28** (helm chart `cloudnative-pg-0.27.0`, appVersion 1.28.0). 1.26+ deprecates the in-tree `barmanObjectStore` in favour of the Barman Cloud Plugin, but the plugin is **not deployed**, and `barmanObjectStore` is still fully functional on 1.28. So this uses the in-tree mechanism. Migrating to the plugin is a follow-up (noted in `docs/cnpg-backups.md`). ## Schedule / retention (defaults — Ben to adjust) | App | Cluster | Bucket | Nightly base backup | | --- | --- | --- | --- | | authentik | postgres | cnpg-authentik | 01:00 | | litellm | litellm-postgres | cnpg-litellm | 01:20 | | artifactapi | postgres | cnpg-artifactapi | 01:40 | | woodpecker | woodpecker-postgres | cnpg-woodpecker | 02:00 | | puppet | puppet-postgres | cnpg-puppet | 02:20 | | encapi | postgres | cnpg-encapi | 02:40 | | paperclip | paperclip-postgres | cnpg-paperclip | 03:00 | | grafana | postgres | cnpg-grafana | 03:20 | Retention is **30d** across the board — flagged as a default to tune per cluster. Schedules are staggered 20 min apart so 8 base backups don't hit RGW at once. ## Validation - `kustomize build --enable-helm` + `kubeconform` (repo CI args, incl. the new ceph schemas) pass on all 8 affected overlays (paperclip validated at base — it has no overlay yet). ceph CRs resolve their schemas (`Skipped: 0`). - `pre-commit run` passes on all changed files (yamllint, no-plain-secrets, etc.). - Note: a full `ci/validate-apps.sh` run aborts locally on the unrelated `cattle-system` overlay (`chart requires kubeVersion < 1.35 vs host helm v1.36.0`) — pre-existing, reproduces on `origin/main`, unrelated to this change. ## Notes / caveats - No overlap with the woodpecker chart-bump PR (#297, overlay files only) or the logging PR (#296) beyond the three **identical** generated `schemas/ceph.unkin.net/*.json` files, which merge cleanly whichever lands first. - Credentials: no manual seeding — the cephrgw-operator mints the RGW user + keys. The only prerequisite is the operator being healthy (it is, in `cephrgw-system`). ## Follow-ups - Barman Cloud Plugin migration (deploy plugin, move clusters to `ObjectStore` CRs). - Tune per-cluster retention / schedule if the defaults don't fit. https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv Reviewed-on: #298 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net>
147 lines
4.9 KiB
Markdown
147 lines
4.9 KiB
Markdown
# CNPG backups
|
|
|
|
Every CNPG Postgres cluster in this repo backs up to a dedicated Ceph RGW (S3)
|
|
bucket: continuous WAL archiving plus a staggered nightly base backup, 30-day
|
|
retention, compressed. Buckets and their S3 keys are provisioned by the in-estate
|
|
`cephrgw-operator` — nothing is seeded by hand. Because the operator's CRs are
|
|
namespace-scoped and CNPG reads its credential Secret from its own namespace, the
|
|
topology is **one bucket per cluster** (`s3://cnpg-<app>`), not one shared bucket.
|
|
|
|
The backup mechanism is CNPG's in-tree `barmanObjectStore` (the deployed operator
|
|
is v1.28; the Barman Cloud Plugin is not deployed — see the migration note at the
|
|
bottom).
|
|
|
|
## Where it lives
|
|
|
|
Per cluster, two files under `apps/base/<app>/`:
|
|
|
|
- `cnpg_cluster.yaml` — the `spec.backup` stanza (WAL archiving + retention).
|
|
- `cnpg_backup.yaml` — the `ObjectStoreUser` + `Bucket` (RGW provisioning) and the
|
|
nightly `ScheduledBackup`.
|
|
|
|
## The backup stanza (`spec.backup` in the Cluster)
|
|
|
|
```yaml
|
|
spec:
|
|
backup:
|
|
retentionPolicy: 30d # default; adjust per cluster
|
|
barmanObjectStore:
|
|
destinationPath: s3://cnpg-<app>
|
|
endpointURL: https://s3.ceph.unkin.net
|
|
endpointCA: # trust the internal Vault PKI CA
|
|
name: vault-ca-cert # reflected into every namespace
|
|
key: ca.crt
|
|
s3Credentials:
|
|
accessKeyId:
|
|
name: cnpg-<app>-backup-s3 # minted by the ObjectStoreUser
|
|
key: AWS_ACCESS_KEY_ID
|
|
secretAccessKey:
|
|
name: cnpg-<app>-backup-s3
|
|
key: AWS_SECRET_ACCESS_KEY
|
|
serverName: <app> # path prefix inside the bucket
|
|
data:
|
|
compression: bzip2 # base backup (zstd not supported here)
|
|
jobs: 2
|
|
wal:
|
|
compression: zstd # WAL segments
|
|
maxParallel: 2
|
|
```
|
|
|
|
Setting `spec.backup.barmanObjectStore` turns on **continuous WAL archiving**
|
|
immediately (CNPG points `archive_command` at the object store). The nightly
|
|
base backup is a separate object:
|
|
|
|
## The nightly base backup (`ScheduledBackup`)
|
|
|
|
```yaml
|
|
apiVersion: postgresql.cnpg.io/v1
|
|
kind: ScheduledBackup
|
|
metadata:
|
|
name: cnpg-<app>-nightly
|
|
namespace: <app>
|
|
spec:
|
|
schedule: "0 0 1 * * *" # 6-field cron, SECONDS first (01:00:00 daily)
|
|
immediate: false
|
|
backupOwnerReference: self
|
|
method: barmanObjectStore
|
|
cluster:
|
|
name: <cluster-name>
|
|
```
|
|
|
|
Schedules are staggered so the base backups don't hit RGW at once:
|
|
|
|
| App | Cluster | Bucket | Nightly (local) |
|
|
| --- | --- | --- | --- |
|
|
| authentik | postgres | cnpg-authentik | 01:00 |
|
|
| litellm | litellm-postgres | cnpg-litellm | 01:20 |
|
|
| artifactapi | postgres | cnpg-artifactapi | 01:40 |
|
|
| woodpecker | woodpecker-postgres | cnpg-woodpecker | 02:00 |
|
|
| puppet | puppet-postgres | cnpg-puppet | 02:20 |
|
|
| encapi | postgres | cnpg-encapi | 02:40 |
|
|
| paperclip | paperclip-postgres | cnpg-paperclip | 03:00 |
|
|
| grafana | postgres | cnpg-grafana | 03:20 |
|
|
|
|
## Where the credentials come from
|
|
|
|
The `ObjectStoreUser` in `cnpg_backup.yaml` tells `cephrgw-operator` to mint an RGW
|
|
user and write its keys into a Secret; the `Bucket` makes that user the bucket owner
|
|
(full read/write on its own bucket). No keys are ever committed.
|
|
|
|
```yaml
|
|
apiVersion: ceph.unkin.net/v1alpha1
|
|
kind: ObjectStoreUser
|
|
metadata:
|
|
name: cnpg-<app>-backup
|
|
namespace: <app>
|
|
spec:
|
|
uid: cnpg-<app>-backup # RGW users are global; keep it unique
|
|
secretName: cnpg-<app>-backup-s3 # -> AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY
|
|
retainOnDelete: true
|
|
---
|
|
apiVersion: ceph.unkin.net/v1alpha1
|
|
kind: Bucket
|
|
metadata:
|
|
name: cnpg-<app>
|
|
namespace: <app>
|
|
spec:
|
|
bucketName: cnpg-<app>
|
|
ownerRef: cnpg-<app>-backup
|
|
retainOnDelete: true
|
|
```
|
|
|
|
Confirm the operator provisioned everything:
|
|
|
|
```bash
|
|
kubectl -n <app> get objectstoreuser,bucket
|
|
kubectl -n <app> get secret cnpg-<app>-backup-s3 \
|
|
-o jsonpath='{.data.AWS_ACCESS_KEY_ID}' | base64 -d; echo
|
|
```
|
|
|
|
## Checking backup status
|
|
|
|
```bash
|
|
# Cluster health + first-recoverability point (WAL archiving working?)
|
|
kubectl -n <app> get cluster <cluster-name> \
|
|
-o jsonpath='{.status.firstRecoverabilityPoint}{"\n"}'
|
|
|
|
# Backups taken so far
|
|
kubectl -n <app> get backups.postgresql.cnpg.io
|
|
|
|
# With the cnpg kubectl plugin (richer view, shows archiving + last backup)
|
|
kubectl cnpg status <cluster-name> -n <app>
|
|
|
|
# Kick a one-off backup right now (verify the whole path end to end)
|
|
kubectl cnpg backup <cluster-name> -n <app>
|
|
|
|
# Look at what actually landed in the bucket
|
|
aws --endpoint-url https://s3.ceph.unkin.net s3 ls s3://cnpg-<app>/<app>/
|
|
```
|
|
|
|
## Follow-up: Barman Cloud Plugin migration
|
|
|
|
CNPG 1.26+ deprecates the in-tree `barmanObjectStore` in favour of the Barman Cloud
|
|
Plugin (removal is planned for a future major). The plugin is **not** deployed today,
|
|
so this repo stays on the in-tree mechanism, which is fully functional on 1.28.
|
|
Migrating means deploying the plugin and moving each cluster's config to an
|
|
`ObjectStore` CR + `plugins:` reference — track that as separate work.
|