9db52c5e264e82d00446e5bec446ebe2384647cd
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ba7a1a9509 |
Enable Reloader secret watching, scope existing auto to configmap-only (#326) (#339)
## Why The re-keyed internal `unkin.net` intermediate broke CA consumers (CNPG->RGW backups, subPath/startup-cached CA mounts) and needed manual pod restarts, because Reloader was deployed with `ignoreSecrets: true` and could not restart on the `vault-ca-cert` Secret. Enabling secret watching naively is unsafe: many workloads carry the generic `reloader.stakater.com/auto`, and the estate rotates numerous Secrets via Vault/VSO — those would restart on every rotation. This enables secret watching but scopes existing `auto` to ConfigMaps, making secret-reload opt-in per Secret. ## Changes - Set `reloader.ignoreSecrets: false` (au-syd1 reloader-system values) so Secrets are watched. - Convert every generic `reloader.stakater.com/auto: "true"` to the ConfigMap-only `configmap.reloader.stakater.com/auto: "true"` — 22 annotations across 19 files. Existing ConfigMap-reload behaviour is preserved; Vault/VSO Secret rotations no longer restart these workloads. - Add explicit `secret.reloader.stakater.com/reload: "vault-ca-cert"` to the CA consumers that mount the CA and carry a Reloader annotation: `artifactapi/api`, `cephrgw-operator`, `puppetserver-master`, `puppetserver-compiler`, `litellm`, `logarchiver`. - Add `secret.reloader.stakater.com/reload: "kanidm-tls"` so kanidm rolls when cert-manager renews its leaf. - Add `docs/ca-rotation.md` runbook (indexed in `docs/README.md`). ## Safety review (secret-only / CA workloads) `vault-ca-cert` is a plain reflected Secret that bootstraps Vault trust (not VSO-rotated; changes only on intermediate re-key). `kanidm-tls` is a cert-manager leaf. Everything else mounted (`environment`, `*-credentials`, `eyaml-keys`, `puppetboard-secrets`, `s3-credentials`, `nats-auth`, `clickhouse-credentials`, `woodpecker-*`) is VSO/CNPG Vault-rotated and deliberately excluded. - `cephrgw-operator` — mounts only Secrets (`cephrgw-credentials` VSO + `vault-ca-cert`), no ConfigMap. Its old comment said "restart when the credentials Secret rotates"; `cephrgw-credentials` is VSO so that is now excluded, and reload is scoped to `vault-ca-cert` only. Comment updated. - `nats` (logging) — old comment "Roll the StatefulSet when nats-auth changes"; `nats-auth` is VSO, so this is now ConfigMap-only (deliberately no roll on rotation). Comment updated. Same for the vector agent/aggregator/vm-ingest (VSO `nats-auth`/`clickhouse-credentials`). - `artifactapi/ui` — mounts neither a ConfigMap nor a Secret; its `auto` was already a no-op. Left as ConfigMap-only. - `puppetdb` / `puppetboard` — mount a ConfigMap plus VSO Secrets (postgres creds / puppetboard-secrets); ConfigMap-only is correct, no secret reload added. CA consumers that mount `vault-ca-cert` but have **no** Reloader annotation (CRD-managed or startup-cached) are documented in `docs/ca-rotation.md` for manual restart rather than annotated here: `grafana`, `observability/vmagent`, `paperclip`, `argocd-repo-server`, plus CNPG clusters (`kubectl cnpg restart`). ## Notes / coordination - Annotations left in their existing location (some sit on the pod template, e.g. `litellm`, `puppetdb`; Reloader reads controller-level metadata — placement unchanged from before, no regression). - Touches `apps/overlays/au-syd1/logging/values-vector-*.yaml`, which overlap open PR #320 (Tier-2 Vector pipelines) — only the one-line reloader annotation is changed here. ## Validation - `make kubeconform` — touched overlays (reloader-system, logging, woodpecker, authentik) valid; only the known-unrelated cattle-system rancher chart kubeVersion failure remains. - `uvx pre-commit run --all-files` — all hooks pass. Closes #326 --------- Co-authored-by: Ben Vincent <neotheo@gmail.com> Reviewed-on: #339 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net> |
||
|
|
a04dcc2975 |
Add k8s Gitea deployment (migration target for git.unkin.net) (#309)
Stand up the git.unkin.net forge on k8s to replace the Puppet VM. Deployed HA-shaped to match what the VM already runs (multi-replica on shared storage + external DB/cache), so this is genuine multi-replica HA rather than single-replica failover. Serves a temporary git2.k8s.syd1.au.unkin.net host; the git.unkin.net cutover is staged in docs/gitea-migration.md. - add apps/base/gitea: namespace, CNPG gitea-postgres (2 instances, S3 backup bucket cnpg-gitea, nightly 04:00/30d), pgbouncer pooler, standalone Valkey (session/cache/queue, AOF), VaultAuth + VaultStaticSecrets, Gateway + HTTPRoute - add apps/overlays/au-syd1/gitea: official Gitea chart 12.6.0 (app 1.26.2, rootless, 2 replicas) via helm-through-kustomize; RWX CephFS repo storage, external CNPG + Valkey, Actions disabled, container registry disabled (moved to artifactapi), Authentik OIDC with auto-register/account-linking; SSH via LoadBalancer VIP 198.18.200.10:2222 - register gitea in the platform ApplicationSet + AppProject - add docs/gitea-migration.md staged cutover plan (VM Postgres->CNPG dump/restore, DNS in main.unkin.net zone, consumer checklist, rollback) Depends on: terraform-authentik gitea OIDC app, and terraform-artifactapi ^gitea/ dockerhub allowlist (both separate PRs). One-time Vault seeds are listed in the migration doc. Reviewed-on: #309 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net> |
||
|
|
dfb495d771 |
Trust internal CA for Authentik SSO; canonical identity.unkin.net for NetBox (#314)
Authentik is canonical at https://identity.unkin.net, served by the internal unkin.net CA. Grafana, LiteLLM and NetBox failed OIDC discovery because their images don't trust that CA (x509: unknown authority); NetBox also still pointed at the secondary admin host. - grafana: mount the reflected vault-ca-cert; set generic_oauth `tls_client_ca`. - litellm: `combine-certs` init builds public+internal CA bundle; `SSL_CERT_FILE` + `REQUESTS_CA_BUNDLE` point at it. - netbox: flip OIDC issuer to identity.unkin.net; same combine bundle for python-social-auth (`requests`). - docs: record the Rancher manual runtime step (issuer + CA in the auth config). Validated: kustomize build + kubeconform + pre-commit. https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv Reviewed-on: #314 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net> |
||
|
|
3c2bdf307a |
Add S3 backups to all CNPG Postgres clusters (#298)
## Why None of the 8 CNPG Postgres clusters in this repo had **any** backup configured. A lost PVC, a fat-fingered migration, or a bad app deploy meant permanent, unrecoverable data loss for authentik, litellm, artifactapi, woodpecker, puppet, encapi, paperclip and grafana. This adds continuous WAL archiving + a nightly base backup to Ceph RGW for every cluster, plus code-forward restore docs. ## What - **`spec.backup.barmanObjectStore`** on each `cnpg_cluster.yaml` — turns on continuous WAL archiving to `s3://cnpg-<app>`, WAL compressed with zstd, base backups with bzip2, 30-day retention. TLS to `s3.ceph.unkin.net` is trusted via the reflected `vault-ca-cert` (`endpointCA`). - **`cnpg_backup.yaml`** per app — a cephrgw `ObjectStoreUser` + `Bucket` (operator provisions the bucket and mints the S3 key into `cnpg-<app>-backup-s3`; **nothing is hardcoded**) and a staggered nightly `ScheduledBackup`. - **`schemas/ceph.unkin.net/*.json`** — the three cephrgw CRD schemas so kubeconform can validate the new CRs. - **`docs/`** — new docs folder (README index + `cnpg-backups.md` + `cnpg-restore.md`). ## Design decisions (answers to the open questions) **One bucket for all, or per-database?** → **Per-database (one bucket + owner user per cluster).** The cephrgw CRDs are namespace-scoped (`BucketRef`/`OwnerRef` resolve only *within the same namespace*), and CNPG reads its S3 credential Secret from its *own* namespace. A single shared bucket would require either cross-namespace bucket refs (unsupported) or hand-copying the S3 secret into all 8 namespaces (defeats "operator mints the keys"). Per-namespace `s3://cnpg-<app>` with a dedicated owner user is the simplest correct topology and needs zero manual seeding. Each user owns exactly one bucket, so owner-level (full) access is already tightly scoped — no extra `BucketAccess` grant needed. **Backup mechanism.** The deployed CNPG operator is **v1.28** (helm chart `cloudnative-pg-0.27.0`, appVersion 1.28.0). 1.26+ deprecates the in-tree `barmanObjectStore` in favour of the Barman Cloud Plugin, but the plugin is **not deployed**, and `barmanObjectStore` is still fully functional on 1.28. So this uses the in-tree mechanism. Migrating to the plugin is a follow-up (noted in `docs/cnpg-backups.md`). ## Schedule / retention (defaults — Ben to adjust) | App | Cluster | Bucket | Nightly base backup | | --- | --- | --- | --- | | authentik | postgres | cnpg-authentik | 01:00 | | litellm | litellm-postgres | cnpg-litellm | 01:20 | | artifactapi | postgres | cnpg-artifactapi | 01:40 | | woodpecker | woodpecker-postgres | cnpg-woodpecker | 02:00 | | puppet | puppet-postgres | cnpg-puppet | 02:20 | | encapi | postgres | cnpg-encapi | 02:40 | | paperclip | paperclip-postgres | cnpg-paperclip | 03:00 | | grafana | postgres | cnpg-grafana | 03:20 | Retention is **30d** across the board — flagged as a default to tune per cluster. Schedules are staggered 20 min apart so 8 base backups don't hit RGW at once. ## Validation - `kustomize build --enable-helm` + `kubeconform` (repo CI args, incl. the new ceph schemas) pass on all 8 affected overlays (paperclip validated at base — it has no overlay yet). ceph CRs resolve their schemas (`Skipped: 0`). - `pre-commit run` passes on all changed files (yamllint, no-plain-secrets, etc.). - Note: a full `ci/validate-apps.sh` run aborts locally on the unrelated `cattle-system` overlay (`chart requires kubeVersion < 1.35 vs host helm v1.36.0`) — pre-existing, reproduces on `origin/main`, unrelated to this change. ## Notes / caveats - No overlap with the woodpecker chart-bump PR (#297, overlay files only) or the logging PR (#296) beyond the three **identical** generated `schemas/ceph.unkin.net/*.json` files, which merge cleanly whichever lands first. - Credentials: no manual seeding — the cephrgw-operator mints the RGW user + keys. The only prerequisite is the operator being healthy (it is, in `cephrgw-system`). ## Follow-ups - Barman Cloud Plugin migration (deploy plugin, move clusters to `ObjectStore` CRs). - Tune per-cluster retention / schedule if the defaults don't fit. https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv Reviewed-on: #298 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net> |