ee2476fe5d461f9426bc2a46efd0c257069eba08
A single Grafana pod with a 1 cpu/1Gi limit is a single point of failure and gets throttled under dashboard load. State lives in Postgres, so extra replicas are safe; spreading them keeps a node loss from taking out Grafana. - set `replicas: 3` on the grafana deployment - raise grafana container limits to 2 cpu / 4Gi - spread grafana pods across nodes with a soft hostname topology constraint Reviewed-on: #508 Co-authored-by: unkin-agent <unkin-agent@unkin.net> Co-committed-by: unkin-agent <unkin-agent@unkin.net>
argocd-apps docs
Operational notes for the manifests in this repo.
| Doc | What it covers |
|---|---|
| cnpg-backups.md | How CNPG Postgres backups (WAL archiving + nightly base backups) to Ceph RGW are configured. |
| cnpg-restore.md | Restoring a CNPG cluster: full recovery, point-in-time recovery, cutover, and gotchas. |
| authentik-rancher-sso.md | Manual runtime step to point Rancher's OIDC auth at the canonical identity.unkin.net issuer and trust the internal CA. |
| gitea-migration.md | Staged cutover of the git.unkin.net forge from the Puppet VM to the gitea namespace. |
| ca-rotation.md | Rolling the internal unkin.net PKI CA (vault-ca-cert): what Reloader restarts automatically vs. manual/CNPG restarts. |
Description
Languages
Shell
91.1%
Makefile
8.9%