None of the CNPG clusters had any backup configured, so a lost PVC or a bad
migration meant permanent data loss. This adds continuous WAL archiving plus a
nightly base backup to Ceph RGW for every cluster, with restore docs.
- Add spec.backup.barmanObjectStore (in-tree; operator is CNPG 1.28, Barman
Cloud Plugin not deployed) to each cnpg_cluster.yaml: WAL zstd, base bzip2,
30-day retention, endpointCA via the reflected vault-ca-cert.
- Add cnpg_backup.yaml per app: a cephrgw ObjectStoreUser + Bucket (one
dedicated s3://cnpg-<app> bucket and owner user per cluster, since cephrgw CRs
and the CNPG credential Secret are namespace-scoped) and a staggered nightly
ScheduledBackup. Credentials are minted by the operator; nothing is hardcoded.
- Add the three ceph.unkin.net CRD schemas so kubeconform can validate the CRs.
- Add docs/ (README index, cnpg-backups.md, cnpg-restore.md) covering config and
full/PITR restore procedures.
Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
## Summary
- Changes `limits.memory` from `1024Mi` to `1Gi` (same value, canonical form)
- Changes `limits.cpu` from `1` (integer) to `"1"` (string, canonical form)
## Why
Kubernetes normalizes resource quantities on write — `1024Mi` becomes `1Gi` and integer `1` becomes string `"1"`. ArgoCD diffs by string comparison, so these equivalent values cause a permanent OutOfSync on the `litellm-postgres` Cluster.
Reviewed-on: #163
finding litellm performance has dropped, crashed in multiple cases, and
then it had scaled to the maximum level using the majority of memory in
cluster.
- reduce the rate at which litellm autoscales
- increase the requests/limits to match usage
Reviewed-on: #144
Deploys LiteLLM proxy with CNPG PostgreSQL (3-instance HA), PgBouncer
pooler, and Redis cache. Introduces a dedicated aitooling AppProject and
ApplicationSet to keep AI tooling services separate from platform infra.
Reviewed-on: #94