grafana: scale to 3 replicas, raise limits #508

Merged
benvin merged 2 commits from benvin/grafana-scale into main 2026-10-02 22:15:26 +10:00
Member

A single Grafana pod with a 1 cpu/1Gi limit is a single point of failure and gets throttled under dashboard load. State lives in Postgres, so extra replicas are safe; spreading them keeps a node loss from taking out Grafana.

  • set replicas: 3 on the grafana deployment
  • raise grafana container limits to 2 cpu / 4Gi
  • spread grafana pods across nodes with a soft hostname topology constraint
A single Grafana pod with a 1 cpu/1Gi limit is a single point of failure and gets throttled under dashboard load. State lives in Postgres, so extra replicas are safe; spreading them keeps a node loss from taking out Grafana. - set `replicas: 3` on the grafana deployment - raise grafana container limits to 2 cpu / 4Gi - spread grafana pods across nodes with a soft hostname topology constraint
unkin-agent added 1 commit 2026-10-02 22:10:15 +10:00
grafana: run 3 replicas, raise limits to 2 cpu/4Gi
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
1c05f1b0a0
Author
Member
  • apps/base/grafana/grafana.yaml:12 — replicas: 3 with no ha_peers/unified_alerting HA config in config means each replica evaluates and fires every alert (duplicate notifications) → set unified_alerting.ha_peers/ha_listen_address to the headless service, or disable HA/alerting evaluation on all but one replica; confirm in the PR what happens to alerting.
  • apps/base/grafana/grafana.yaml:12 — no podAntiAffinity/topologySpreadConstraints, so 3 replicas can land on one node and the SPOF claim is not fixed → add a topologySpreadConstraint on kubernetes.io/hostname under deployment.spec.template.spec.
  • apps/base/grafana/grafana.yaml:35 — limits go to 2 cpu/4Gi while requests stay 100m/256Mi, 3 replicas can burst to 6 cpu/12Gi and are first to be evicted/throttled unevenly → raise requests to realistic values, or justify the 4Gi.
  • nit: PR mixes two concerns (HA replicas, resource limits) → fine as one PR only if the body states the limits bump is needed for the same reason; otherwise split into "scale to 3 replicas" and "raise limits".
- apps/base/grafana/grafana.yaml:12 — replicas: 3 with no ha_peers/unified_alerting HA config in `config` means each replica evaluates and fires every alert (duplicate notifications) → set `unified_alerting.ha_peers`/`ha_listen_address` to the headless service, or disable HA/alerting evaluation on all but one replica; confirm in the PR what happens to alerting. - apps/base/grafana/grafana.yaml:12 — no podAntiAffinity/topologySpreadConstraints, so 3 replicas can land on one node and the SPOF claim is not fixed → add a topologySpreadConstraint on kubernetes.io/hostname under `deployment.spec.template.spec`. - apps/base/grafana/grafana.yaml:35 — limits go to 2 cpu/4Gi while requests stay 100m/256Mi, 3 replicas can burst to 6 cpu/12Gi and are first to be evicted/throttled unevenly → raise requests to realistic values, or justify the 4Gi. - nit: PR mixes two concerns (HA replicas, resource limits) → fine as one PR only if the body states the limits bump is needed for the same reason; otherwise split into "scale to 3 replicas" and "raise limits".
unkin-agent added 1 commit 2026-10-02 22:12:06 +10:00
grafana: spread replicas across nodes
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
d7c776bc45
Author
Member

No findings.

No findings.
benvin merged commit ee2476fe5d into main 2026-10-02 22:15:26 +10:00
benvin deleted branch benvin/grafana-scale 2026-10-02 22:15:26 +10:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: unkin/argocd-apps#508