Compare commits

..

2 Commits

Author SHA1 Message Date
Ben Vincent ffcf646d87 Rework jellyfin-ha into a true-HA StatefulSet deployment
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline failed
Turn the single-replica jellyfin-ha app into a proper high-availability
deployment so the fork's Redis-coordinated distributed transcoding and
PostgreSQL main database can actually be exercised.

- Replace the Deployment with a 2-replica StatefulSet for stable pod
  identity; set JELLYFIN_INSTANCE_ID from metadata.name (the fork's Redis
  transcode-lease owner id), add soft podAntiAffinity and a PDB
  minAvailable 1.
- Move the main Jellyfin DB to PostgreSQL via a CloudNativePG trio
  (3-instance Cluster, PgBouncer Pooler, Ceph RGW barman backups) mirroring
  the litellm pattern; an init container writes database.xml selecting the
  fork's Jellyfin-PostgreSQL provider and the DSN is composed from the
  CNPG-generated app secret pointed at the pooler.
- Share /config on an RWX cephfs PVC across replicas; keep /cache per-pod
  via a volumeClaimTemplate.
- Fix the transcode mount to the fork's real path /config/transcodes on the
  RWX PVC (raid5) so a surviving pod can resume the segments of the pod it
  takes over.
- Add Intel iGPU hardware transcoding via the gpu.intel.com/i915 device
  plugin resource plus render/video supplemental groups.
- Switch the Service to sessionAffinity ClientIP to reduce transcode churn.
- Disable UDP auto-discovery. Library scans still run on every replica; a
  single-scanner leader election is a planned follow-up.
2026-08-10 23:27:36 +10:00
Ben Vin a29fa5b8cc Add jellyfin (HA fork) app under a new media project
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Deploys the jellyfin-ha fork (git.unkin.net/unkin/jellyfin-ha) to au-syd1
via ArgoCD, under a dedicated media AppProject/ApplicationSet rather than
extending platform.

- New media AppProject + media-apps ApplicationSet (watches
  apps/overlays/*/jellyfin); registered in the argocd kustomizations
- apps/base/jellyfin: namespace, deployment (single replica to start),
  service, in-namespace Redis (transcode session store), gateway + httproute
  at jellyfin.k8s.syd1.au.unkin.net
- Storage: RWO config (cephrbd), RWX transcode scratch and RWX media
  library (cephfs) per the HA fork's pod-takeover requirement
- au-syd1 overlay
2026-07-05 22:32:24 +10:00
210 changed files with 697 additions and 6679 deletions
-29
View File
@@ -1,29 +0,0 @@
when:
- event: pull_request
steps:
- name: vector-test
image: artifactapi.k8s.syd1.au.unkin.net/dockerhub/timberio/vector:0.57.0-debian
commands:
# Dummy creds + writable dirs so the full topologies build; the unit tests
# only exercise the transforms (sources are not started).
- export CLICKHOUSE_USER=ci CLICKHOUSE_PASSWORD=ci
- export NATS_PRODUCER_PASSWORD=ci NATS_CONSUMER_PASSWORD=ci
- mkdir -p /vector-data-dir /etc/vault-ca
- cp /etc/ssl/certs/ca-certificates.crt /etc/vault-ca/ca.crt
# Transform tier + VM ingest: unit-tested transforms.
- vector test apps/base/logging/vector/aggregator.yaml apps/base/logging/vector/aggregator-tests.yaml
- vector test apps/base/logging/vector/vm-ingest.yaml apps/base/logging/vector/vm-ingest-tests.yaml
# Agent + archiver have no transforms to unit-test; validate they build.
- vector validate --no-environment apps/base/logging/vector/agent.yaml
- vector validate --no-environment apps/base/logging/vector/archiver.yaml
backend_options:
kubernetes:
serviceAccountName: default
resources:
requests:
memory: 256Mi
cpu: 250m
limits:
memory: 1Gi
cpu: 1
-1
View File
@@ -8,7 +8,6 @@ resources:
- httproute.yaml
- namespace.yaml
- service.yaml
- vpa.yaml
configMapGenerator:
- name: age-api-config
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: age-api-vpa
namespace: age-api
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: age-api
updatePolicy:
updateMode: "Off"
+1 -1
View File
@@ -35,7 +35,7 @@ spec:
mountPath: /combined-certs
containers:
- name: api
image: git.unkin.net/unkin/artifactapi:v3.7.7
image: git.unkin.net/unkin/artifactapi:v3.7.6
imagePullPolicy: IfNotPresent
ports:
- containerPort: 8000
-55
View File
@@ -1,55 +0,0 @@
---
# Ceph RGW (S3) backup target for the artifactapi CNPG cluster, provisioned by the
# in-estate cephrgw-operator. One dedicated bucket + owner user per cluster:
# cephrgw CRs are namespace-scoped and CNPG reads its S3 credential Secret from
# its own namespace, so backups are per-database rather than one shared bucket.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-artifactapi-backup
namespace: artifactapi
spec:
displayName: "CNPG backup owner (artifactapi)"
# RGW users are global; keep the uid namespace-qualified so it never collides.
uid: cnpg-artifactapi-backup
maxBuckets: 5
# Operator writes AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (+ RGW_UID,
# S3_ENDPOINT) into this Secret; the Cluster's barmanObjectStore consumes it.
secretName: cnpg-artifactapi-backup-s3
# Keep the RGW user (and thus the keys) if this CR is ever deleted, so an
# in-flight restore can still reach the archive.
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-artifactapi
namespace: artifactapi
spec:
bucketName: cnpg-artifactapi
# The owner user has full control of its own bucket (read + write), which is
# all the backup/restore identity needs — no extra BucketAccess grant.
ownerRef: cnpg-artifactapi-backup
versioning: false
tags:
app: artifactapi
purpose: cnpg-backup
# Never drop the backups if the CR is removed; retire buckets by hand.
retainOnDelete: true
---
# Nightly base backup. Continuous WAL archiving is always-on via the Cluster's
# spec.backup.barmanObjectStore; this schedules the periodic full backup that
# WAL is layered on top of. Schedules are staggered across clusters so the 8
# base backups do not hit RGW at once (CNPG cron is 6-field, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-artifactapi-nightly
namespace: artifactapi
spec:
schedule: "0 40 1 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: postgres
+1 -30
View File
@@ -7,35 +7,6 @@ metadata:
spec:
affinity:
podAntiAffinityType: preferred
backup:
# 30-day retention (DEFAULT — adjust per cluster if needed). Enforced by CNPG
# against the object store on each successful base backup.
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-artifactapi
endpointURL: https://s3.ceph.unkin.net
# radosgw serves a Vault-PKI cert; trust the internal CA (reflected into
# every namespace as the vault-ca-cert Secret).
endpointCA:
name: vault-ca-cert
key: ca.crt
# Keys minted by the ObjectStoreUser in cnpg_backup.yaml; never hardcoded.
s3Credentials:
accessKeyId:
name: cnpg-artifactapi-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-artifactapi-backup-s3
key: AWS_SECRET_ACCESS_KEY
# Path prefix within the bucket; keep stable across restores (see docs).
serverName: artifactapi
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: artifacts
@@ -108,7 +79,7 @@ spec:
cpu: 500m
memory: 512Mi
requests:
cpu: 50m
cpu: 250m
memory: 256Mi
smartShutdownTimeout: 180
startDelay: 3600
-2
View File
@@ -7,7 +7,6 @@ resources:
- api-hpa.yaml
- configmap.yaml
- cnpg_cluster.yaml
- cnpg_backup.yaml
- cnpg_pooler.yaml
- gateway.yaml
- httproute.yaml
@@ -18,4 +17,3 @@ resources:
- ui-hpa.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
- vpa.yaml
+1 -1
View File
@@ -22,7 +22,7 @@ spec:
automountServiceAccountToken: true
containers:
- name: ui
image: git.unkin.net/unkin/artifactapi-ui:v3.7.7
image: git.unkin.net/unkin/artifactapi-ui:v3.7.6
imagePullPolicy: IfNotPresent
ports:
- containerPort: 80
-45
View File
@@ -1,45 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: api-vpa
namespace: artifactapi
# NOTE: this workload also has an HPA. updateMode Off is recommendation-only
# and does not act, so there is no HPA/VPA conflict today. Do not flip to Auto/
# Initial without first moving the HPA off CPU/memory (VPA owns those under Auto).
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: api
updatePolicy:
updateMode: "Off"
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: redis-vpa
namespace: artifactapi
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: redis
updatePolicy:
updateMode: "Off"
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: ui-vpa
namespace: artifactapi
# NOTE: this workload also has an HPA. updateMode Off is recommendation-only
# and does not act, so there is no HPA/VPA conflict today. Do not flip to Auto/
# Initial without first moving the HPA off CPU/memory (VPA owns those under Auto).
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: ui
updatePolicy:
updateMode: "Off"
-55
View File
@@ -1,55 +0,0 @@
---
# Ceph RGW (S3) backup target for the authentik CNPG cluster, provisioned by the
# in-estate cephrgw-operator. One dedicated bucket + owner user per cluster:
# cephrgw CRs are namespace-scoped and CNPG reads its S3 credential Secret from
# its own namespace, so backups are per-database rather than one shared bucket.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-authentik-backup
namespace: authentik
spec:
displayName: "CNPG backup owner (authentik)"
# RGW users are global; keep the uid namespace-qualified so it never collides.
uid: cnpg-authentik-backup
maxBuckets: 5
# Operator writes AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (+ RGW_UID,
# S3_ENDPOINT) into this Secret; the Cluster's barmanObjectStore consumes it.
secretName: cnpg-authentik-backup-s3
# Keep the RGW user (and thus the keys) if this CR is ever deleted, so an
# in-flight restore can still reach the archive.
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-authentik
namespace: authentik
spec:
bucketName: cnpg-authentik
# The owner user has full control of its own bucket (read + write), which is
# all the backup/restore identity needs — no extra BucketAccess grant.
ownerRef: cnpg-authentik-backup
versioning: false
tags:
app: authentik
purpose: cnpg-backup
# Never drop the backups if the CR is removed; retire buckets by hand.
retainOnDelete: true
---
# Nightly base backup. Continuous WAL archiving is always-on via the Cluster's
# spec.backup.barmanObjectStore; this schedules the periodic full backup that
# WAL is layered on top of. Schedules are staggered across clusters so the 8
# base backups do not hit RGW at once (CNPG cron is 6-field, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-authentik-nightly
namespace: authentik
spec:
schedule: "0 0 1 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: postgres
+3 -34
View File
@@ -7,35 +7,6 @@ metadata:
spec:
affinity:
podAntiAffinityType: preferred
backup:
# 30-day retention (DEFAULT — adjust per cluster if needed). Enforced by CNPG
# against the object store on each successful base backup.
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-authentik
endpointURL: https://s3.ceph.unkin.net
# radosgw serves a Vault-PKI cert; trust the internal CA (reflected into
# every namespace as the vault-ca-cert Secret).
endpointCA:
name: vault-ca-cert
key: ca.crt
# Keys minted by the ObjectStoreUser in cnpg_backup.yaml; never hardcoded.
s3Credentials:
accessKeyId:
name: cnpg-authentik-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-authentik-backup-s3
key: AWS_SECRET_ACCESS_KEY
# Path prefix within the bucket; keep stable across restores (see docs).
serverName: authentik
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: authentik
@@ -106,12 +77,10 @@ spec:
resources:
limits:
cpu: 500m
# 512Mi OOMKilled replicas under load (shared_buffers 128MB +
# max_connections 200 leave no headroom) — see incident 2026-07-28.
memory: 1Gi
requests:
cpu: 50m
memory: 512Mi
requests:
cpu: 250m
memory: 256Mi
smartShutdownTimeout: 180
startDelay: 3600
stopDelay: 1800
-2
View File
@@ -4,7 +4,6 @@ kind: Kustomization
resources:
- cnpg_cluster.yaml
- cnpg_backup.yaml
- cnpg_pooler.yaml
- gateway.yaml
- httproute.yaml
@@ -18,4 +17,3 @@ resources:
- redis-service.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
- vpa.yaml
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: redis-vpa
namespace: authentik
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: redis
updatePolicy:
updateMode: "Off"
@@ -24,6 +24,3 @@ spec:
- 198.18.27.0/24
- 198.18.28.0/24
- 198.18.29.0/24
# Admin/management access (individual hosts, not whole subnets)
- 10.10.12.200/32 # benvin workstation (wireguard)
- 198.18.21.160/32 # benvin router
@@ -13,11 +13,9 @@ spec:
storageSize: 2Gi
# Restrict queries to internal networks (puppet acl-main.unkin.net).
# 10.42.0.0/16 (pod net) is required so secondaries can SOA-refresh
# from the primary during catalog replication. localhost is required so the
# operator's in-pod `nsupdate` (sent to 127.0.0.1) passes query-authorization;
# without it every dynamic update is "denied due to allow-query".
# from the primary during catalog replication.
extraOptions:
- "allow-query { localhost; auth-acl-main; 10.42.0.0/16; }"
- "allow-query { auth-acl-main; 10.42.0.0/16; }"
service:
type: LoadBalancer
externalTrafficPolicy: Local
@@ -33,7 +31,7 @@ spec:
external-dns.alpha.kubernetes.io/hostname: bind-authoritative-primary.k8s.syd1.au.unkin.net
resources:
requests:
cpu: 20m
cpu: 100m
memory: 128Mi
limits:
cpu: "1"
@@ -6,5 +6,4 @@ resources:
- cluster.yaml
- tsigkey.yaml
- zones.yaml
- records.yaml
- acls.yaml
@@ -1,64 +0,0 @@
# Individually-managed authoritative records for the unkin.net zone.
# DNSRecords must live in the same namespace as their BindZone (the operator
# resolves zoneRef/clusterRef/updateKeyRef within the record's namespace), so
# these sit alongside the zone in bind-internal, not in the app namespace.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
# "internal" in the name distinguishes this from the external DNS that
# Authentik will manage its own records from later.
name: identity-dns-internal
namespace: bind-internal
spec:
zoneRef: unkin-net
name: identity
type: A
ttl: 600
values:
# traefik-internal gateway VIP; the authentik Gateway serves the
# identity.unkin.net hostname there.
- 198.18.200.4
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: s3-ceph-cname
namespace: bind-internal
spec:
zoneRef: ceph-unkin-net
name: s3
type: CNAME
ttl: 600
values:
# radosgw S3 endpoint. Points at the Consul service for now; the real
# target will be changed later.
- radosgw.service.consul.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: dashboard-ceph-cname
namespace: bind-internal
spec:
zoneRef: ceph-unkin-net
name: dashboard
type: CNAME
ttl: 600
values:
# Ceph mgr dashboard, reached via lb1. Lets in-cluster clients (the
# cephrgw-operator) resolve dashboard.ceph.unkin.net.
- lb1.unkin.net.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: lb1-unkin-net
namespace: bind-internal
spec:
zoneRef: unkin-net
name: lb1
type: A
ttl: 600
values:
- 103.216.191.185
@@ -9,16 +9,3 @@ metadata:
spec:
clusterRef: bind-authoritative
algorithm: hmac-sha256
---
# Client-update key: puppet clients (profiles::dns::updater) nsupdate their own
# records to the authoritative zones with this key. Operator generates the
# material into Secret client-update-tsig; the same value must reach puppet
# eyaml (or the planned Vault-sync bridge) for clients to authenticate.
apiVersion: bind.unkin.net/v1alpha1
kind: BindTSIGKey
metadata:
name: client-update
namespace: bind-internal
spec:
clusterRef: bind-authoritative
algorithm: hmac-sha256
@@ -15,8 +15,6 @@ spec:
zoneName: unkin.net
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -28,24 +26,6 @@ spec:
zoneName: main.unkin.net
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
# ceph.unkin.net: the ceph host (ausyd1nxvm2069/halb) publishes
# dashboard.ceph.unkin.net via nsupdate; puppet targets a dedicated
# `zone ceph.unkin.net.`, so it must exist here or the update gets NOTZONE.
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
metadata:
name: ceph-unkin-net
namespace: bind-internal
spec:
clusterRef: bind-authoritative
zoneName: ceph.unkin.net
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -57,8 +37,6 @@ spec:
zoneName: 13.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -70,8 +48,6 @@ spec:
zoneName: 14.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -83,8 +59,6 @@ spec:
zoneName: 15.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -96,8 +70,6 @@ spec:
zoneName: 16.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -109,8 +81,6 @@ spec:
zoneName: 17.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -122,8 +92,6 @@ spec:
zoneName: 19.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -135,8 +103,6 @@ spec:
zoneName: 20.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -148,8 +114,6 @@ spec:
zoneName: 21.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -161,8 +125,6 @@ spec:
zoneName: 22.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -174,8 +136,6 @@ spec:
zoneName: 23.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -187,8 +147,6 @@ spec:
zoneName: 24.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -200,8 +158,6 @@ spec:
zoneName: 25.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -213,8 +169,6 @@ spec:
zoneName: 26.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -226,8 +180,6 @@ spec:
zoneName: 27.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -239,8 +191,6 @@ spec:
zoneName: 28.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
@@ -252,5 +202,3 @@ spec:
zoneName: 29.18.198.in-addr.arpa
type: primary
defaultTTL: 600
dynamicUpdate: true
updateKeyRef: client-update
@@ -23,7 +23,7 @@ spec:
type: ClusterIP
resources:
requests:
cpu: 20m
cpu: 100m
memory: 128Mi
limits:
cpu: "1"
@@ -1,10 +1,6 @@
---
# Key that external-dns (and DNSRecord objects) use to send RFC2136 dynamic
# updates to the primary. The operator generates the material into a Secret
# (externaldns-key-tsig) in this namespace. secretTemplate stamps emberstack
# reflector hints onto that Secret so it is mirrored into the externaldns
# namespace, where the external-dns controller reads it -- guaranteeing
# external-dns presents exactly the key the primary's allow-update accepts.
# updates to the primary. The operator generates the material into a Secret.
apiVersion: bind.unkin.net/v1alpha1
kind: BindTSIGKey
metadata:
@@ -13,9 +9,3 @@ metadata:
spec:
clusterRef: bind-externaldns
algorithm: hmac-sha256
secretTemplate:
annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "externaldns"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "externaldns"
@@ -7,4 +7,3 @@ resources:
- authoritative
- resolvers
- externaldns
- tsig-api
@@ -8,15 +8,11 @@ metadata:
spec:
clusterRef: bind-resolvers
entries:
- 10.42.0.0/16 # k8s pod network (kube-proxy masquerades node-originated LB queries)
- 198.18.1.10/32
- 198.18.2.160/27
- 198.18.21.160/27
- 198.18.2.192/27
- 198.18.21.192/27
# Admin/management access
- 10.10.12.200/32 # benvin workstation (wireguard)
- 198.18.21.160/32 # benvin router (also within 198.18.21.160/27 above)
- 198.18.13.0/24
- 198.18.14.0/24
- 198.18.15.0/24
+1 -10
View File
@@ -21,18 +21,9 @@ spec:
forwarders:
- 8.8.8.8
- 1.1.1.1
# The internal split-horizon zones are served UNSIGNED by the in-cluster
# authoritative, but their public parents publish DS records (e.g. unkin.net
# is DNSSEC-signed on the Internet). With dnssec-validation on, the validator
# sees "parent indicates secure" but gets an insecure answer and returns
# SERVFAIL (broken trust chain). Treat the forwarded internal domains as
# insecure so they are not validated. unkin.net covers all *.unkin.net
# (incl. k8s.syd1.au.unkin.net); 18.198.in-addr.arpa covers every reverse zone.
extraOptions:
- "validate-except { unkin.net; 18.198.in-addr.arpa; consul; }"
resources:
requests:
cpu: 20m
cpu: 100m
memory: 128Mi
limits:
cpu: "1"
@@ -1,9 +1,6 @@
# Conditional forward zones, from the puppet openforwarder view.
# Upstreams: unkin authoritative 198.18.200.6, consul 198.18.19.14,
# k8s 198.18.200.8 (in-cluster bind-externaldns VIP).
# k8s -> in-cluster bind-externaldns 198.18.200.8 for both the forward zone
# k8s.syd1.au.unkin.net and the reverse zone 200.18.198.in-addr.arpa, which
# external-dns now publishes to (see the external-dns migration PRs).
# Upstreams: unkin authoritative 198.18.200.6, consul 198.18.19.14, k8s 198.18.200.8.
# k8s -> in-cluster bind-externaldns 198.18.200.8.
# (Zones that forwarded to 10.10.16.x were dropped; consul left as-is.)
---
apiVersion: bind.unkin.net/v1alpha1
@@ -64,22 +61,6 @@ spec:
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
metadata:
name: fwd-200-18-198-in-addr-arpa
namespace: bind-internal
spec:
clusterRef: bind-resolvers
viewRef: openforwarder
zoneName: 200.18.198.in-addr.arpa
type: forward
catalog: false
forwarders:
# Reverse zone for the k8s LB range, published by external-dns to the
# in-cluster bind-externaldns alongside k8s.syd1.au.unkin.net.
- 198.18.200.8
---
apiVersion: bind.unkin.net/v1alpha1
kind: BindZone
metadata:
name: fwd-13-18-198-in-addr-arpa
namespace: bind-internal
@@ -1,27 +0,0 @@
---
# Companion TSIG API. The operator reconciles this into a Deployment, Service,
# ConfigMap, master-token Secret and namespaced RBAC. vault-plugin-secrets-bind-tsig
# calls it to create/rotate/delete TSIG keys, which it does by managing
# BindTSIGKey resources in this namespace (the operator reconciles the material).
#
# The master access token Secret (bind-tsig-api-token) is generated by the
# operator if absent; a VaultStaticSecret may later pre-seed/overwrite it so the
# token is sourced from Vault rather than generated in-cluster.
apiVersion: bind.unkin.net/v1alpha1
kind: BindTSIGAPI
metadata:
name: bind-tsig-api
namespace: bind-internal
spec:
image: git.unkin.net/unkin/bind-tsig-api:v0.2.3
replicas: 1
port: 8443
# targetNamespace defaults to this resource's namespace (bind-internal), where
# the authoritative cluster and its keys live.
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 250m
memory: 128Mi
+1 -1
View File
@@ -21,7 +21,7 @@ spec:
runAsNonRoot: true
containers:
- name: operator
image: git.unkin.net/unkin/bind-operator:v0.2.6
image: git.unkin.net/unkin/bind-operator:v0.1.5
args:
- --metrics-bind-address=:8080
- --health-probe-bind-address=:8081
+1 -2
View File
@@ -6,7 +6,6 @@ resources:
- namespace.yaml
# CRDs are pulled from the bind-operator repo at the matching tag rather than
# vendored here, so they never drift from the operator.
- https://git.unkin.net/unkin/bind-operator/raw/tag/v0.2.6/config/crd/install.yaml
- https://git.unkin.net/unkin/bind-operator/raw/tag/v0.1.5/config/crd/install.yaml
- rbac.yaml
- deployment.yaml
- vpa.yaml
+1 -9
View File
@@ -23,15 +23,7 @@ rules:
resources: ["pods/exec"]
verbs: ["create", "get"]
- apiGroups: ["apps"]
resources: ["statefulsets", "deployments"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
# The BindTSIGAPI reconciler deploys the companion API: a ServiceAccount plus
# a namespaced Role/RoleBinding granting it access to BindTSIGKey + Secrets.
- apiGroups: [""]
resources: ["serviceaccounts"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
- apiGroups: ["rbac.authorization.k8s.io"]
resources: ["roles", "rolebindings"]
resources: ["statefulsets"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
- apiGroups: [""]
resources: ["events"]
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: bind-operator-vpa
namespace: bind-system
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: bind-operator
updatePolicy:
updateMode: "Off"
-82
View File
@@ -1,82 +0,0 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: cephrgw-operator
namespace: cephrgw-system
labels:
app.kubernetes.io/name: cephrgw-operator
annotations:
# Restart the operator when the credentials Secret rotates.
reloader.stakater.com/auto: "true"
spec:
replicas: 1
selector:
matchLabels:
app.kubernetes.io/name: cephrgw-operator
template:
metadata:
labels:
app.kubernetes.io/name: cephrgw-operator
spec:
serviceAccountName: cephrgw-operator
securityContext:
runAsNonRoot: true
containers:
- name: operator
image: git.unkin.net/unkin/cephrgw-operator:v0.3.1
args:
- --metrics-bind-address=:8080
- --health-probe-bind-address=:8081
- --leader-elect
envFrom:
# Provides CEPH_RGW_ACCESS_KEY/SECRET_KEY and the endpoints
# (CEPH_RGW_ENDPOINT / CEPH_RGW_ADMIN_ENDPOINT), plus optional
# CEPH_RGW_REGION / CEPH_RGW_CA / CEPH_RGW_INSECURE. Rendered from
# Vault per docs/ceph-setup.md; not managed in GitOps.
- secretRef:
name: cephrgw-credentials
env:
# Trust the internal unkin.net (Vault PKI) CA so the operator can
# verify radosgw's TLS cert. vault-ca-cert is reflected into every
# namespace from the certificates namespace.
- name: CEPH_RGW_CA_FILE
value: /etc/vault-ca/ca.crt
volumeMounts:
- name: vault-ca-cert
mountPath: /etc/vault-ca/ca.crt
subPath: ca.crt
readOnly: true
ports:
- containerPort: 8080
name: metrics
- containerPort: 8081
name: health
readinessProbe:
httpGet:
path: /readyz
port: 8081
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /healthz
port: 8081
initialDelaySeconds: 15
periodSeconds: 20
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
volumes:
- name: vault-ca-cert
secret:
secretName: vault-ca-cert
@@ -1,14 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
# CRDs are pulled from the cephrgw-operator repo at the matching tag rather
# than vendored here, so they never drift from the operator.
- https://git.unkin.net/unkin/cephrgw-operator/raw/tag/v0.3.1/config/crd/install.yaml
- rbac.yaml
- deployment.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
- vpa.yaml
-5
View File
@@ -1,5 +0,0 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: cephrgw-system
-42
View File
@@ -1,42 +0,0 @@
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: cephrgw-operator
namespace: cephrgw-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: cephrgw-operator
rules:
- apiGroups: ["ceph.unkin.net"]
resources: ["*"]
verbs: ["*"]
# The operator delivers RGW access/secret keys into Secrets.
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
- apiGroups: [""]
resources: ["events"]
verbs: ["create", "patch"]
- apiGroups: ["coordination.k8s.io"]
resources: ["leases"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
# v0.3.1 startup check reads its own CRDs to warn if they are stale/missing.
- apiGroups: ["apiextensions.k8s.io"]
resources: ["customresourcedefinitions"]
verbs: ["get", "list"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: cephrgw-operator
subjects:
- kind: ServiceAccount
name: cephrgw-operator
namespace: cephrgw-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: cephrgw-operator
-21
View File
@@ -1,21 +0,0 @@
---
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultAuth
metadata:
name: default
namespace: cephrgw-system
spec:
method: kubernetes
mount: k8s/au/syd1
vaultConnectionRef: vso-system/default
allowedNamespaces:
- cephrgw-system
kubernetes:
# Shared "default" role: binds the namespace's default ServiceAccount and
# grants the templated kv/kubernetes/namespace/<ns>/<sa>/* read policy, so
# no per-app terraform-vault change is needed.
role: default
serviceAccount: default
audiences:
- vault
tokenExpirationSeconds: 600
@@ -1,30 +0,0 @@
---
# Renders the radosgw credentials from Vault into the cephrgw-credentials
# Secret the operator Deployment consumes via envFrom. The KV secret's keys
# (CEPH_RGW_ACCESS_KEY/SECRET_KEY, CEPH_RGW_ENDPOINT, optional
# CEPH_RGW_ADMIN_ENDPOINT/REGION/CA) are copied verbatim, so they land as the
# matching env vars.
#
# The path sits under the templated default policy
# (kv/data/kubernetes/namespace/<ns>/<sa>/*), so it needs no dedicated Vault
# role or policy. Seed the values with:
# vault kv put kv/kubernetes/namespace/cephrgw-system/default/cephrgw-credentials \
# CEPH_RGW_ENDPOINT=https://s3.ceph.unkin.net \
# CEPH_RGW_ADMIN_ENDPOINT=https://radosgw.service.consul:443 \
# CEPH_RGW_ACCESS_KEY=... CEPH_RGW_SECRET_KEY=...
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultStaticSecret
metadata:
name: cephrgw-credentials
namespace: cephrgw-system
spec:
vaultAuthRef: default
mount: kv
type: kv-v2
path: kubernetes/namespace/cephrgw-system/default/cephrgw-credentials
refreshAfter: 5m
hmacSecretData: true
destination:
name: cephrgw-credentials
create: true
overwrite: true
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: cephrgw-operator-vpa
namespace: cephrgw-system
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: cephrgw-operator
updatePolicy:
updateMode: "Off"
@@ -7,4 +7,3 @@ resources:
- serviceaccount.yaml
- clusterrole.yaml
- clusterrolebinding.yaml
- vmservicescrape.yaml
@@ -1,17 +0,0 @@
---
# Scrape cert-manager webhook + cainjector metrics endpoints.
# Picked up by the observability VMAgent (selectAllByDefault).
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: cert-manager
namespace: cert-manager
spec:
selector:
matchLabels:
app.kubernetes.io/instance: cert-manager
endpoints:
- port: metrics
path: /metrics
- port: http-metrics
path: /metrics
@@ -1,6 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
@@ -1,7 +0,0 @@
---
apiVersion: v1
kind: Namespace
metadata:
labels:
app.kubernetes.io/name: clickhouse-system
name: clickhouse-system
-30
View File
@@ -1,30 +0,0 @@
# consul (k8s)
Consul server cluster (DC `au-syd1`), deployed via the HashiCorp helm chart with
ACLs enabled (`default_policy: deny`, parity with the VM cluster).
## API access (ACL auth)
The HTTP API and UI are served on port 8500 behind the gateway at
`https://consul.k8s.syd1.au.unkin.net` (and `https://consul.service.consul`).
With ACLs enabled, requests beyond the anonymous policy require a token:
```bash
# management (bootstrap) token — seeded from Vault, synced by VSO into the
# consul-bootstrap-acl-token secret; same value as the VM cluster's
# initial_management token:
CONSUL_HTTP_TOKEN=$(vault kv get -field=token kv/kubernetes/namespace/consul/default/bootstrap-acl-token)
curl -H "X-Consul-Token: $CONSUL_HTTP_TOKEN" https://consul.k8s.syd1.au.unkin.net/v1/status/leader
# consul CLI:
CONSUL_HTTP_ADDR=https://consul.k8s.syd1.au.unkin.net CONSUL_HTTP_TOKEN=$CONSUL_HTTP_TOKEN consul members
```
The UI at the same hostname exposes an ACL login (top right) — paste a token.
Anonymous requests get the anonymous-token policy only (reads for DNS/service
discovery; no writes, no ACL/token APIs).
Prefer short-lived tokens minted by Vault's consul secrets engine over the
management token for day-to-day use; the terraform-* CI roles already work this
way.
+4 -4
View File
@@ -46,8 +46,8 @@ spec:
- backendRefs:
- group: ""
kind: Service
name: consul-http
port: 8500
name: consul-ui
port: 80
weight: 1
matches:
- path:
@@ -74,8 +74,8 @@ spec:
- backendRefs:
- group: ""
kind: Service
name: consul-http
port: 8500
name: consul-ui
port: 80
weight: 1
matches:
- path:
-3
View File
@@ -6,6 +6,3 @@ resources:
- namespace.yaml
- gateway.yaml
- httproute.yaml
- service.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
-25
View File
@@ -1,25 +0,0 @@
---
# ClusterIP service targeting the consul server pods' HTTP API (8500).
# The HashiCorp chart only ships consul-ui (also 8500 via the server pods)
# and the headless consul-server; this named service gives the Gateway a
# stable API backend. Consul serves both the HTTP API and the UI (at /ui/)
# on this same port, so routing the API hostname here preserves the UI too.
apiVersion: v1
kind: Service
metadata:
name: consul-http
namespace: consul
labels:
app.kubernetes.io/name: consul
app.kubernetes.io/instance: consul
spec:
type: ClusterIP
selector:
app: consul
component: server
release: consul
ports:
- name: http
port: 8500
protocol: TCP
targetPort: 8500
-18
View File
@@ -1,18 +0,0 @@
---
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultAuth
metadata:
name: default
namespace: consul
spec:
allowedNamespaces:
- consul
kubernetes:
audiences:
- vault
role: default
serviceAccount: default
tokenExpirationSeconds: 600
method: kubernetes
mount: k8s/au/syd1
vaultConnectionRef: vso-system/default
-17
View File
@@ -1,17 +0,0 @@
---
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultStaticSecret
metadata:
name: bootstrap-acl-token
namespace: consul
spec:
destination:
create: true
name: consul-bootstrap-acl-token
overwrite: true
hmacSecretData: true
mount: kv
path: kubernetes/namespace/consul/default/bootstrap-acl-token
refreshAfter: 5m
type: kv-v2
vaultAuthRef: default
-1
View File
@@ -7,4 +7,3 @@ resources:
- vaultauth.yaml
- vaultstaticsecret.yaml
- storageclass.yaml
- vmservicescrape.yaml
-15
View File
@@ -1,15 +0,0 @@
---
# Scrape the ceph-csi-cephfs nodeplugin + provisioner http-metrics endpoints.
# Picked up by the observability VMAgent (selectAllByDefault).
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: ceph-csi-cephfs
namespace: csi-cephfs
spec:
selector:
matchLabels:
app: ceph-csi-cephfs
endpoints:
- port: http-metrics
path: /metrics
-1
View File
@@ -7,4 +7,3 @@ resources:
- vaultauth.yaml
- vaultstaticsecret.yaml
- storageclass.yaml
- vmservicescrape.yaml
@@ -1,15 +0,0 @@
---
# Scrape the ceph-csi-rbd nodeplugin + provisioner http-metrics endpoints.
# Picked up by the observability VMAgent (selectAllByDefault).
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: ceph-csi-rbd
namespace: csi-cephrbd
spec:
selector:
matchLabels:
app: ceph-csi-rbd
endpoints:
- port: http-metrics
path: /metrics
-55
View File
@@ -1,55 +0,0 @@
---
# Ceph RGW (S3) backup target for the encapi CNPG cluster, provisioned by the
# in-estate cephrgw-operator. One dedicated bucket + owner user per cluster:
# cephrgw CRs are namespace-scoped and CNPG reads its S3 credential Secret from
# its own namespace, so backups are per-database rather than one shared bucket.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-encapi-backup
namespace: encapi
spec:
displayName: "CNPG backup owner (encapi)"
# RGW users are global; keep the uid namespace-qualified so it never collides.
uid: cnpg-encapi-backup
maxBuckets: 5
# Operator writes AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (+ RGW_UID,
# S3_ENDPOINT) into this Secret; the Cluster's barmanObjectStore consumes it.
secretName: cnpg-encapi-backup-s3
# Keep the RGW user (and thus the keys) if this CR is ever deleted, so an
# in-flight restore can still reach the archive.
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-encapi
namespace: encapi
spec:
bucketName: cnpg-encapi
# The owner user has full control of its own bucket (read + write), which is
# all the backup/restore identity needs — no extra BucketAccess grant.
ownerRef: cnpg-encapi-backup
versioning: false
tags:
app: encapi
purpose: cnpg-backup
# Never drop the backups if the CR is removed; retire buckets by hand.
retainOnDelete: true
---
# Nightly base backup. Continuous WAL archiving is always-on via the Cluster's
# spec.backup.barmanObjectStore; this schedules the periodic full backup that
# WAL is layered on top of. Schedules are staggered across clusters so the 8
# base backups do not hit RGW at once (CNPG cron is 6-field, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-encapi-nightly
namespace: encapi
spec:
schedule: "0 40 2 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: postgres
+1 -30
View File
@@ -7,35 +7,6 @@ metadata:
spec:
affinity:
podAntiAffinityType: preferred
backup:
# 30-day retention (DEFAULT — adjust per cluster if needed). Enforced by CNPG
# against the object store on each successful base backup.
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-encapi
endpointURL: https://s3.ceph.unkin.net
# radosgw serves a Vault-PKI cert; trust the internal CA (reflected into
# every namespace as the vault-ca-cert Secret).
endpointCA:
name: vault-ca-cert
key: ca.crt
# Keys minted by the ObjectStoreUser in cnpg_backup.yaml; never hardcoded.
s3Credentials:
accessKeyId:
name: cnpg-encapi-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-encapi-backup-s3
key: AWS_SECRET_ACCESS_KEY
# Path prefix within the bucket; keep stable across restores (see docs).
serverName: encapi
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: encapi
@@ -108,7 +79,7 @@ spec:
cpu: 500m
memory: 512Mi
requests:
cpu: 50m
cpu: 250m
memory: 256Mi
smartShutdownTimeout: 180
startDelay: 3600
-2
View File
@@ -10,8 +10,6 @@ resources:
- gateway.yaml
- httproute.yaml
- cnpg_cluster.yaml
- cnpg_backup.yaml
- cnpg_pooler.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
- vpa.yaml
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: encapi-vpa
namespace: encapi
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: encapi
updatePolicy:
updateMode: "Off"
-55
View File
@@ -1,55 +0,0 @@
---
# Ceph RGW (S3) backup target for the grafana CNPG cluster, provisioned by the
# in-estate cephrgw-operator. One dedicated bucket + owner user per cluster:
# cephrgw CRs are namespace-scoped and CNPG reads its S3 credential Secret from
# its own namespace, so backups are per-database rather than one shared bucket.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-grafana-backup
namespace: grafana
spec:
displayName: "CNPG backup owner (grafana)"
# RGW users are global; keep the uid namespace-qualified so it never collides.
uid: cnpg-grafana-backup
maxBuckets: 5
# Operator writes AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (+ RGW_UID,
# S3_ENDPOINT) into this Secret; the Cluster's barmanObjectStore consumes it.
secretName: cnpg-grafana-backup-s3
# Keep the RGW user (and thus the keys) if this CR is ever deleted, so an
# in-flight restore can still reach the archive.
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-grafana
namespace: grafana
spec:
bucketName: cnpg-grafana
# The owner user has full control of its own bucket (read + write), which is
# all the backup/restore identity needs — no extra BucketAccess grant.
ownerRef: cnpg-grafana-backup
versioning: false
tags:
app: grafana
purpose: cnpg-backup
# Never drop the backups if the CR is removed; retire buckets by hand.
retainOnDelete: true
---
# Nightly base backup. Continuous WAL archiving is always-on via the Cluster's
# spec.backup.barmanObjectStore; this schedules the periodic full backup that
# WAL is layered on top of. Schedules are staggered across clusters so the 8
# base backups do not hit RGW at once (CNPG cron is 6-field, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-grafana-nightly
namespace: grafana
spec:
schedule: "0 20 3 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: postgres
-87
View File
@@ -1,87 +0,0 @@
---
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres
namespace: grafana
spec:
affinity:
podAntiAffinityType: preferred
backup:
# 30-day retention (DEFAULT — adjust per cluster if needed). Enforced by CNPG
# against the object store on each successful base backup.
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-grafana
endpointURL: https://s3.ceph.unkin.net
# radosgw serves a Vault-PKI cert; trust the internal CA (reflected into
# every namespace as the vault-ca-cert Secret).
endpointCA:
name: vault-ca-cert
key: ca.crt
# Keys minted by the ObjectStoreUser in cnpg_backup.yaml; never hardcoded.
s3Credentials:
accessKeyId:
name: cnpg-grafana-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-grafana-backup-s3
key: AWS_SECRET_ACCESS_KEY
# Path prefix within the bucket; keep stable across restores (see docs).
serverName: grafana
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: grafana
encoding: UTF8
localeCType: C
localeCollate: C
owner: grafana
secret:
name: postgres-credentials
enablePDB: true
enableSuperuserAccess: false
failoverDelay: 0
imageName: ghcr.io/cloudnative-pg/postgresql:18.1-system-trixie
instances: 2
logLevel: info
monitoring:
customQueriesConfigMap:
- key: queries
name: cnpg-default-monitoring
disableDefaultQueries: false
enablePodMonitor: false
postgresql:
parameters:
max_connections: "200"
shared_buffers: 128MB
primaryUpdateMethod: restart
primaryUpdateStrategy: unsupervised
replicationSlots:
highAvailability:
enabled: true
slotPrefix: _cnpg_
synchronizeReplicas:
enabled: true
updateInterval: 30
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 50m
memory: 256Mi
smartShutdownTimeout: 180
startDelay: 3600
stopDelay: 1800
storage:
resizeInUseVolumes: true
size: 10Gi
storageClass: cephrbd-fast-delete
switchoverDelay: 3600
@@ -1,13 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDashboard
metadata:
name: bind9-exporter-dns
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
resyncPeriod: 5m
allowCrossNamespaceImport: false
gzipJson: H4sIAAycS2oC/+1da1PbuBr+Kx5v5wx0A8QJ4ZKZ/QCBnu2Z0naBdmany8kotpKodSwfSaZkMzm/fd9Xsh2HXEhCYiDkC8SWLOm9PM+rq92zSRBwRRTjgbSrPdtnUtnVbz27ETFfvQ/sqlOwPaKI5JFwKWZR3RD+Z28W7Ih5cKslSJMExO4XbBqQhg/ZlIhowW4zL/3NXB7UuM8FPCBaDbJVLFglx4E/lUrBcrahtIB0sIaTQdOsf1knPhVKQqoiokWVaWyHwQ+nWCzYHaLc9knQtatN4kuK2Vog0bebwqDBst3gRHjYvtF7N3DXo9IVLMQqIfGUBd6xdfbxyrqi4pa51LrC5kjFXLkLDaEeU1kpm0y6xP+TEgH5hLrggWrbVWhbK6DqPSjIKZWLxwVUU9i+5txXLNTpPgt+SK11VA7UTO8UFQHxLUyxh2RhCiu0z5KGW1eJDVRSom2n8sXPR8KHi7ZSoazu7bWYakeNXZd39kLqEqkE34tNt5MqRO41fN7Y60AyFXuh4B2q2jSC26iVHXoXcgEpO14gd79LaHX/BgW5pR/5z9QGIQmobyRzue+TUFIvTRzrVYOKUq96k8nYR+0x7zPXztrW3gnVlfYL9p1WJdi/CJnwQec404CbxG9MaxasW9Am2tE+0dImtrjqgo46A6UL0EC/sHAlTUZ9r8aDJmvhcx5tkshXWmDX4KZndzgCylZtQWWb+560tfe6rENQXAfxEIYsaBlxeZgCXOMEHg0i37dRIgmF6/aB08H9j3sg3AAgEjyEEd/IO6ht0ATSkNyPFDogaCFMja3x3RKUBpByS/wIMmOdqJgkWYA3pIlHRQ3BKEBI21ogfksFWJtqA46Yfl+b/iBjeSex/ABSN6iJuzNQ8GfOAtSiZouMQnRjLowwAQ9QEA3Q4VvfI4B9sxvfJJHicJMLRgPDUAgvuPwbME+MXr3IpZ8y1RDfxQbZHUoC+yY2szRo1SqQMTJQ+doWvfhXtlK4FfpRiwVfqZCm3uPdyq5jL8W/AdVolze9el3xKqTRqqRASZ7sWzsWKHIL0V9vcK7qinVoPU7tsUAqErj0t///Zb8JoL3V3bd/2f1tC2oyaSrl/wBoAwQ2kidX74ir0CfQdLRFA+8dF+CqJpMgQYumvDEAoXE58AE0aRIW7CFsXlKJXKwdLXFpsNgGn/ni82AIn6UNPleKT1c752wIfXqAmtZGYixGic+I1N1FaRTeIEJmuhCy/YEGLexnOUVzTeVjexhTAJ667RjvbzJAbLVsfvxbEA+dz/T/xuDCKWlg4L8YGdBJ9mgAXU2WEQHbV05UjgWQ21aa6EZC6Dria3CAwW8WpL+BFAZ9MsURAcnVfc8GAfHSeAJe/GQearc8BFkkCo3XBISGrDKoJdhbv04IyRSI2KCgY4BgxlsmwSXE8lGLYKhqJb6WGScLPCoowqHpc4WupjX3KWsVuBkSl2ZdBDzL/THQDLhnSL0PIGhm4LAsnLLAFZRIugVPuFTKuhtGCSDr2g4TYFmwvvPGb3AH/gFIvzmlorzZth/CZmksNkfQ6IyCMRM5zCijQy9pyxgzO+z4CCMzz6p9/mJdQ5bsoKMHXkYQw7HvSBgeaOfXHlaPVQfsxG6ZF2HESsGu6RzKuiN3LBu6lKnCeK8pNXFXDbguuaPGTM1EXrSIz1unoHRNVYNnkUwG2drYuolZb3TZLHZk1kqQ1H+RjOTMyEhHmZ50ZRofOQdDhIQKOpHXZiQcZ1uEpQRrtdWVnikY4i1j+cVoy1k+bZk6noa1YmitlLS2ZNTZEtC12EI6mpmyEpLathpdawsRvG29taAzt23tWTgvQsSWy6NATSoWHwHKi0AhSICT+ivIgu2sk8xLib2ehVVZ/f7SyPF5cWDsrPfoTaPOBnvYMersor0APxrAzUGQ9gdOPMupdKDYX2q1k7Jugbmp7x0dF98VS8m9OJ9TqTmH9tqR69EM5Hr4Ysi1tF7kmkuXUJOfD57ujPKoJrgFOnkJmGags8LS2l+ZFAcWFaCSFeB09QI4S5fAGRKhtkYhZVw82ASRpQWRylxBxDmYIYocveYu+pPOLOTRR9ck1qEdLrr1C9q5RsXXG11FJ87zwfTg1r2n3sHk7wMP/WplnzmNmk1Q5lzP1Ijbpt70RxaZ2Pgi9XRhTgFvHqUtIAyUa2HB+QXAeSy6gEBxubNEw6XKM4u3LSCOKTYjzdkKY3tpVbE95ucJIV7r7AlD/FqM9Q6zU/tmJu14WpjeX87M/jMbxK0+/MbOg9RJ7qx3zKfWWbx1h2vaMQYr9p96uJdM/IP96s3sKlxCSW9wGw3Mdw3N9y8SQsBB8giHiUA8pMFqJfoU6qXpBwPiPMsXYzzluY+0yqUcVjFwmwRACwknM5b6yoSKtMF+KZ7slw9rg3vWhQ64maQXwN+l+fg7XpqdSuCVpyfwzcrsI2js1rhz0n+813FcNqUlgMqTqEWM7JxETIlkycwd882z4+tx3ebNovPsfeWhnbvOwTSuPX56rnU2XPuIXTB6b9r/Igos5EWhz1xYZB5ZVJ5KSKMbYSYz0VlaxVi+3V8u3Y4TEkwAeF2ZgLCQrWsYXsk+zVNCQQF76LGrk/IyrWLsbNL+whHlj8hQzItYbFlBSMnvQEapnOz7dXI6kfE+gMMtsLt69EzG65h0Ku1PC6RTtm8ZD33k0lAMntcZbXv9J14bGuZpFiNBEzZ7dLidY5MVtnuZm6zuQduamb9Lz3ixfHAkZAeBPncRr3qCPZ6gmU52hxuyW9tl8AlUJyiQnVQr7Xnz0F3yJtL7/HZppLA+6ZpeCM9tpj4WnPooTz0BdLQhsVdCYqCKEJRDV0pe5lDpJPJafEh9GTceuEtqvLxWzspvbF0+SsbW5ZzG1mBb7gN8XtHYeljjUxcEncpmlnotWFj7eErHOczl3jL6EygZTkuNn9jNaep6VHQfznzmI/iHk4vz/KeyRyXuMKlfw5CP1Bfvry5Ormu/zzC/vXLJAWoBrpl4+Yh+ffnlY+3k+vxs7N7KR3VEtGBW2iM51/jdDKLWLTRnDy3vT10+dkqbyLxGkXk589mzR+QVTWenTPWiprM3DDUvQx3NwlDlDUOtGUN1n9XAYdl7DKcLDm+GLOYq/jm8qrOYyxmqqYILqvKLTJdQWXdpp44zW2O6my7zWgek+BTxAxFpfxOR1mDP5T2C8iKh34mYvoCnEbk/qMohOvmr6EIbtjrtWmexXBvKWuvNbZXyVMoqbijrRVLWMFO5eOS6LoSEwuc/rfJ8hvX66Lh1eQmvkFevd6/xK9mJ9gAzORtmWqPhPQxvof9UB80zz3Sn8hzrfiU+jM+ecpQ/Kr+MXH3kMK+5jnHbWJZy2BDmEa7Oa6jjWLTX2aO80WpFnUEpFXQTCcGsQ1KQl49Q2V3ffElFZL5UoqM5XHr62A5epF8MQU4YuB5qnXZCH9QctIY+QJOSIiic+tRVmd008bvwP2ddOH5NffZmP3lVaBHdy/Ujj574A9L0SYP6979k0wGXYmmW+Es0QzlSekSX0SOq+2hKtebg7xa9M9vm5Q8WfhH+VTdwB5KMfExH7x9KZe8vHP8g4DF4V79hY3s2TfyHNyap4LtOGid7L1WCLqZuXNewRRTqV1lvZxA6sM8OpOx8JYLhTsYdPYyLtx3Nqj2DPUe73Vdd7R9xU2JXzF4axZm2FvA9s9faV+OQ+XRK/51LVZ2kdmTOBfWeEPAE5WPJM2p/b+vbf6s3v27DC5T21tkSnw1RjjVEaNLmN8SY91Bt7b7dhtj3oImwzllNpC2kS55uovISTIQxSccsEL0JzdV++nNHv/4SvoGhr+w4T8hgbkl/EyVudj0J9DpO6LDiFPFvWf/V7zHVJZnXaZaL+jeGzRL+cTw7Dt31gS2yD+isBzqrfqC0r/9iNDr0dIGenQjwN348BIaesINRmh2NcQ/AfPrrPA5Z+A2w1M+Gv3+FgScJh7g9k9If+uNfus/xD1uJBfNebQAA
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -1,13 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDashboard
metadata:
name: frr-ospf-route-metrics
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
resyncPeriod: 5m
allowCrossNamespaceImport: false
gzipJson: H4sIAAycS2oC/+1ZbW/bNhD+KwJRFC0gF7brZkGAfXDcZgiQpkFcbNiawqCkk82FJjWScuIK/u+7o14sN0lXB0nRJfliS5TIe+655046sWBcKe24E1pZtlcwKaxje58KFuVCukPF9nohS7jjVucmBrrFLTP8Z1PDU644C1kuEjzvdILfyqGg02GrkIHikcQ7nckhZDORNMci1mqkpTY4zUwj/qIbBv1eD3/evAmD3ktcU/E5GRmu0QXPg6EE4yxerSAk3M4izU3CVp/JYCJc22QqbMzln8DN2HHj3mvlZmyvGxL0bPZRa+lE5gekUOfo/qfPIcu4Amk9BakAmYy0SsWU/E4g5bl0nqY5zzKhptUcNzNgZ1om5TVNnjIeWS1zBwjXOsjKJePK66kBUHhlwWWON3dX4fqagWR9ZbeLrqFvegHGIIXe4opcEMmJ9vbQp9chu2B7g5BdeneWfkVGYcHo6ayJrjfxvsRXGqjIqMaUVjT0d26dSJfVYAzKgcFhbQQe+miQf7nTOJiBoRtGM66mMGqtbx1XCcUmJI/yGD60cHAZkydMcpTb57CkGkdY7Tkep1xaQDeQ2YuTtpXqEvIOl66yVqHZPhLEdpv9ryLT88q6QOaP+FLnrtQWDmUynwr1Oxhb0tHrv+q96hIEbqbgSiPtxGHPWmco1suMzKXGTNC0hNhpM8mzojn59Yxpm6VnLBSKuIwBR57Vx2ds5YlNDyn3hpgAaFk4Ej/7MD45CEb1OgGK3+WtrMEFnHf5/6nvwYa+d7bSNzfAr+i7ks49qftYu+NcyrsU+f2q0eC6MHFYduUkFVHBU4HCE9ligFJcmJRUiH+30OXhyWIQnNLyNhAqODjcfzCi3N0Q5S9PovxRoty5A1HuPFBRlpXydSXKwZMo70+U91MlH1h5rJX45kmJP0KJd1kat1NipZpGdhmX4Bx0YiTXipiawxjDqed0D78Udl+bBMwY+WwopOGR7zwg+QuM3rzQiipxzcrhIx6BLGNFpyeSxzDHNdZhiLgZSjFV5WjXD/whEjc74PSujEOv8I0yMfxi7JaeBGwNiexUSPkh47FwKORe2UMmJL7Nzom63ANTOiZhCipZa6LuN6vzhfiylhKGA1tbEp1tbiDLh0RApmUtcRrj1IvRgQfO9vooJS2UG4sv4JOLul54i428EVFeTmy69mqBWr00j9StYOFbPIvt7yYKjHd8jmWF1piizjIvk7AOrfd71S43FXFN8HWaMqoWP1l52vXlqd9t9cw7VX3qb9SnOozrGoHwE2Ezyetq5D+aUEKv9RZpR/ouaT6qlqgSv1FC4fVC6l7zXVFhkSrpidDGNTzfbdWg9nKiQEzxI4qxxXfWipKPA23mnIAVhUjR7dV/tKPHtZkgWgZe1TRrXVacmAOmgMBoPaLikpX591V5+amqS+8OqkvFz6MuLr1BVV1eP6bqQi9vE2n5ROXzCEzx3WXlI73LBEfjYVAUtMjX9SW8JQxwF9qcbw3nuJx3E6D9WwPiNjJboxmO909vgjK6PTf2FnE6HiOMG6C83RYKocCCC0Zhmx/rnMrT90N5V80kOC0U764+jAjviJa3Tw+gpwfQQ3wA7Vx9APX71QNo95E8gGyMXxFgkuTGy2tiAbcdE1ts8WrbbMZ84/X24PS0vdnijQZvK6PX1xfaZjQgNV+nD65NkiGPu5SPNp7BnDc8DHrEgJcXGcQbqJDhH7XnpF8CBPOMMqlUc7ORG+fG+KAVPhFjmScwxI8bteVqp3WDxX9yMEsqD5jr4Gbgt48aiL3WBmwzyQu2MVXWLYwzEC9BTfNaxozS6htxvBloa60apqTKOCk/z7y42tmEjf2X17tRrnO9B1XlbyGvR27rADZY22DH278Nm7Kf1JWJ+BxMGWc6/0I5halq9IX1zX1bsf45+Lz8uhO8B6ylsW228hFGh2B0TP3xZ1ELsb/6F7MoRVUzIAAA
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDashboard
metadata:
name: gitea
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
resyncPeriod: 5m
allowCrossNamespaceImport: false
gzipJson: H4sIAAycS2oC/+2abW/bNhDHv4pBDMMGOIWlxG5qYC+yDimCpknQtOuLNjAo6SwzoUiNpJK4gb/77qgH20G2dbCgtIBe2BAlinf83+9OpKAHxpXSjjuhlWXTByaFdWz6+YH99NOC28VbWLIp09E1xG4ajl6O2ZBFhZDuRLFpMGQJd9zqwsSA3fb2Bm8Mn3PFB3t72BEUjyRecKaAIVuIpDkWsVavtdQG7zJpxH8ZDQdhEODfeDwcBL/izYpnNObR2r3Bz4MjCcZZvOqWOV1N0MVIc5Ow1dUKvQEbG5FTb7z4TivhtBm8EQ74wIK5BUNeJcJt+pUqcCcJzmY/eBVi0/B88UFr6UTOpqMhKqJuUJrPV0OWcwXSenm4FNz6KZBsaDvidDTn0sLQ+3UKKnULHHdUtmHz8oZqCbhwOR4nPD64vonRwbmQ0otLByhoIkBhTMbkm0gutI8TDjwZsjs2PRyye+8nBmq08jInoC7BiA2DAue3jyNKSEEldD+/TZurcWGMN1G1M36/PhaqObYLfdc0HIZFNq1bLovG3sqLRs1SYWrciYTEQD9VIeWFFsq90wSEb+Osdd4wSPKc1qrjYDmgUsrxFBp7uSxSof4EY8tYT168fLGPo+Q0MElW4M1h1V7rgNNMwABhN5fa4Q3WC3WOaKC2UMXZ5jyGzfhZx+ObtQ4O8hwSdHHtkOMmBVeyAfc5WciNjsHamQGLQys3yyDTZjmLlg7sw7WOfvvCUmLzC1uhJ+goIkqaMlZH6libjLvyjIE5YcqOkHW0t8BhF1omlctOZHBsdMampGfZfg9pKWnd4XIh5m7dw1ESsPeVe4PSPUqumv4HDDk3kNSBtNo4H0Mf7lmVhEIl4lYkBbqOwUL1FepXAlXeVvXzmYXD3/N74eMcFfGN16x0KCt5IEeb/C+vlOCVg9WkeTKW/B5Kzee1UtjXUGAlj0DWA0id/s4t+KzydFcWCe7HJlbDVke78l6WE8aika7TybdO4bYeeEWWv7O6cvgNdSXs68qz1JU4L2YW8Ema2JlX7amaghY2IvBcNeb1xcdBldd9cemLS1Vcgsk3VJf9vro8S3XROajZPHliodLwS8k0K135z1pTt4557GjZH3RZfs5xLoPjPy778vOs5effK8h2nTjwdWKysbmZrMqKEG5lXUzVrMpJL5+vRSCT840+XMakJM7UujPtzih3r2i7OOeFdL5PxvNcqLQmaQO9hzp4PLJaFo4sUC6VsYmrXWzCzc1eVNov/SjnfYUzZ3o7UR+XHQ9PXVe0oiGuC+vEfFmdpDri967aUFnl1QaXF04Tlv9UVp7Ibp/GMx5TildZ+79WBE+k4JMZd1ZkEZiBng+O1raqXMEy4dguPEy2eAgOeyB2BcJAri29LBHdQPF+215bYAThdqU46MnYlYwCn++dIPGxMtQaC4dbLIx7FHZFQVhbdFMeTmpLu8NwsLGCCEb1IyP4fmlYgpS4wPohVhGlOp0sImpTbRFx8IiIUU/E7kT4nUQnQJzWltri4XCbh7CvEC3woE3KlfjKO6sT548MtkVHtbBcl4tJj8fueOAv6wSLD5Wh1nCYbONw2NOwOw1zTZ52wsNxY6otIsJHK8ywX0+0gEReRFLge9RlJ1RceGtvS2st7zwO6gfHy++Xi7wwufwhdqIYlW5eSlxWhlredTQ0HPQ07E7DHUQLrW86AeLT2lbLO4+GiVc9E2280JbAbVcvsxtbbe83aij6p0YLTMQ6y0C5Tph4vbbV9qajZmK/Z2J3JjIhwTqtuqkU7zattb31qLkIxz0YLSwquIsX3VDxqTG1hcSVH5pkxlNBRqJit4w3kwhD0nkpoVLXT6cJEGS5RBVUuvFVP31W4j9BwVNz7ySKebcXLPyHMWWLVX1y2ngZ6ll5MauF8KTsj8hh79aY/gL/vz/yxzReSH9BwmqbXylqUxYZ3Hj7QNVC+M/zsV0QvEyd/RWc3I3ffPWQGEFf6ttHU7itFRiv/gYQwxSUwDAAAA==
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDashboard
metadata:
name: nzbget
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
resyncPeriod: 5m
allowCrossNamespaceImport: false
gzipJson: H4sIAAycS2oC/+1X227bRhD9FWIfCqegDcmNAtdAH2wHLgzErgEZDVpbMJbkiNx4uUvsDq0blG/vzHJFybHbpg3aIkBfJHEvczlz5gy1EtIYixKVNV4cr4RWHsXx7UpkrdJ4YcTxMBWFROlt63LgI7ho6FuUTk6lkSIVrSroeX8/+bFbSvb3xToVYGSm6SS6FlJRqaL/rXJrzqy2jq65MpN7gzQ5HA7pYzRKk+ErsmlkzU5OttEl3yQnGhx62o0hFNJXmZWuEOsJOywU7rqcKp9L/QtIN0bp8NIarMTxIOXQm+rGWo2qCQlqZR4o/dtJKhppQPsAwYtpN87WgBW0vs+8ADxcjEaFzF9/eMg59akCXZxZM1UlXy1gKluNAeC8y3slalsEe1IDIuznWnqvwu289WhrPiPnyp9aV4AbV3YmjqdSe0qMl8/AIDgofgVnn26wg8vOOMIcRbf8TmagaSk+XmuZQ002aEm2aGk5k+5Eq9J0q4Ow8F4VWJ3LHDnmwcEb4oKTszEuGGWGDQRnq/VPjcwVLjbwFoqMxCCMDaeYAOeuy0tDCabow8ZNKeLzo1rG34SGMp6qftVqrko8wI4vOP/G6sCOGIx0otsMcYfSNlYZHKslRTJKBRMC3hLHncra7mJf2WiAXHoC+5rveQ4fHoHNemLG0yg8yvxBmVDi0tm2YcLSyXonb7KGlQOyqAsfcetrb6dTsd7Q5VKZSznvjdeyach2ZOXWxs51mXmrW2RwPULTkTbyiwICMLTzKHVLhw1Fvk6328Sc7ebRIPRPaxTT4ZQs0ZOltB3VrDNbS8wrCMQNnM8WV9yhdKyJ2iFuqFN1clMREmXVtMhGqFsaqp6KVsLVLoTe+YrSn0Nx9kncMcewxyBtLofeOOBadWju2GE5IDfDQTocTDpaRpEgA5OQYkkZXdsAIh09SgU11eHrVMwDcZm+zLgiMKfPbMtXwk/qnHMR1K54ZQMvOVqQIWgVPqmM5L9QvtFysWnFoEyEyE7jOVVW3J/Mt3fRBUvXeqcl+mp7YkOw4K3Dnl58VLoS8Esli6LlCG9IMd6TPPY8ZFHt9STnL1qbN1wq5STCnllm5P7ewMzfU6cSa+6RmXCfLRD8qlv64eOdOPj2TqxvR/WE9X1KuF0CyjG1XF713pTJdVvAVdzldPqhYajfGLaNBgTAzq0jalIwq+hpvSbrTpqyHwMOphdFbM3Wwyk17VZ8mFh/H7SXwYGaiuy2CHZo+bbe+xzEAkKv7ozYzMwemj/M/1n3/Q4KpzQriTMKg4J3oSRv3yXjBqBIskVy7ewj+XXbMYuqBoqTm/jL4PoqxyIvj+0USZ83U/H/MfmfjMl/fSR+MgU/a4AcxQFy+CcD5NmACG/eT+dDZpEZ/xUNCP7X0GlHVL2odVwsWdzntjX4T6j/SY7qERLSFgN5h/pfmgLPJfFFiy8I4iSYZirSxjAUi16UavkzOB8a57vvuQI9a8NwwNgO8d8WE4tthpeh0PtUvNn+m4pd2u5JxDONosBZKePzkqtMwDs7851oP8mjLyjA4GFp1AezHJSDnIm+CZC6fwbwEP4jsRKufwNsYW6gFQ4AAA==
File diff suppressed because one or more lines are too long
@@ -1,13 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDashboard
metadata:
name: puppet-report
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
resyncPeriod: 5m
allowCrossNamespaceImport: false
gzipJson: H4sIAAycS2oC/+2be2/bNhDAv4rAbUACOJ0k27FjYAOaFzagj60J9sfaIqAlWiYiiyopOU4977Pvjg9J9pLGbtM2bVy0qsXn8e7440lnzwnNMlHQgotMkcGcpFwVZPD67aJFYqYiyXOsIgNyzArKUxZ7PBsJOdE9PDoUZeH9UeY5KzzJciELb8IKySNFWoTFvKDDlJFBIUvWIknGit9jMuj5fouMecyORFZIkSrXIOXZpcLZW0SykWRqTAYjmiqokuIKa+YkEmlKc8WqmtVZxownY1gDCbp+PgMxcpqx1Ham0Zid8wkDsckgK9O0hQMKeUijy0SKMourcXX5XzQt2XIRjkRkMqQ7Xb/lBb2w5XW6Lc9/ctDbhdl0TdjuQVV4AFU+VvUP6ips3O2Yf9Bpl8BqY1pQJUoZwVyEsndpnvFeycJphxE0RMQnFJdgJF5dMZNSyEpIYx0YB02Q0DJhaNcJndm1BP19UP+EZ/YebtRYXLnBijHqXaTxMzrUarPDVuXPqbxk0lUsGgqH+TjoL4D/soLJKU2dxFy9YFc3WHlC85xnyfl1znS/xr0xWEYnqJIpiuoVwivYrIBppnYpi1bVRNIsuaFJuNDTzI5Bw38IkAvGDdD/UDBd8FzE2D8SWcaigsXE1J3jMFb8XKhixGdmifbmFHz3jL/Hrj3/JywHl63a6M+NJl3dRMv4HLxXr20kxQSqcAqo02IPyIufn+KNcBUovgIPhpXoD/IS1KctOuJpeoQOiYtHz2oH4HNBHy79A3StoI9ON8JRnBGxb6OP6RKiGx+0sbFxBGfYgsqEFUZYNsuxUwF7Z2d3L9c7/sLs+PkYNPLLG/Ij/v+GtFg25VJkE5bp0sbtG7L4GZTve796ewExDKhkcz5zSqMCBYT1pixhWXzqHJpoLCBBCCpJFSyHZh0fdVS5p8Kd3wcLB/2+j6rkBe4UgrveUzyLmJdS5WiFDcAqqFiQr3Z27ZAEmicpU4BH51ANk/a1SXVpZVIBApFfVq1pXbE2qC54Yfw2KqWeGl35Y/Fk9tUSnXaQO/B3HSTdwLE7mbQmgzKRsVsw5N9MoVXc3IkhnH9hyNN9bOTpfmfksba8ATxLxLlgU9gyahPwtNBYUHoOwU4KHHIIups68zl2XSygsQltbpFmfTq1giaXUCBPjLwTN8rjgs9HxEMPkj29LXu+CntW/Cfw70RPs4f/pOt/NH4sUk7huaiUDBizSSD0YABkxSfbqOcbJc/+ljzbqOfDUc9ZGUVMqYeEHSsSqOu7IA/dqViy/tugGlffJnn6W/I8DvKAD2jPrGIe+xj1bUY8r9xq7pE8/jbi+XLcCcIteB4HeKIxrmejmGdxM2JuAYmdYH2SNDhyVPX9gvELnSbGYqDHcyOITdxYqZ7h222bkntlXnJr5qydN2v7/8ub0ZRTdWS5M7cLR2X90Ds5DPePofWpzg5i0UnYOe6EUPQScoPwYktdZ6j3H46OnsLQUP4KVSQL29rvdUMc4AyYGJd2jJ5/fHiiSy85rESXdftBxz+CMn34YEk/POwe92F5ZEhllUX8JKBxdOXA5icR0S+n0GIpT4EAClaBow8w1A2oKsmeqvPmHGiziuf2HHBah6FPJnlx3Sz4m0nh7gEsVV/AaJ0QRWudacFMw2YST4HVjyF9W/mG0Zjtqt2pJiruWLW0f694XIw1qxsgXaWZhUrOQNFZQZNaQzm2kjTmpdLIzS0UneSgKSYZbvZRKnCrKCY5U07Tdj6DpH3cjHAqViuDnQn+8EwTyp0VN+SpRH6502st73VMXs03jVi8YJc0TxuyaQRzA1dWqYLLReFODa7DCTH3Z2M+qk4It7+PS2mS7zspLViVydrVTBdpwXN9ACuMctLSZO+totQY0lx1aIFeog2kHeLCQoxnMZ/yuITFLiqyJZLmY5hhRmdcI2BinMBYDs8zTScjadMTna+9rhwfv1pwTWf29F3KVqcYA7hRUpEcUuUO7Zkr1ntgdR4EXD3S2CT27GhosA+NpUO3xlBvcbBHCDzk2sFnwJptUHPNwWwLNgM2txvXI9ttD2af8iR2P1TzP4w1R7FbsFY9lHnDaw9DqFLdF+H8B0K4ZS59bcq9rXUPO96DL1htGCeG3W81TrQruwdu2nDQ/3RuNjC5DjhtcIimNR/vHZvVk7SjZvD1qGk9y6XLyTI7F/cMVR0n6hdd/74h8B3FEU8AQfD4yEBT/9zHq6+PAG71/Po54FvHlAK07mG7TWhbO9knhpM48VqoNQ68DSUfIhNtLNl5DEz0HxITqzOpIWB4Nyo3erB+VOHndwHDBxd13vrmcumlpfIq7Zv3nYC4Cf0LUgFa6UHbafxICgWfuKw9+VoPF8ObbLQbTbSWzOiY8SjYJIcnC3hZ2/jFA26n1CYpbEbGYWi+uIuL5ovMCL4sSsuYPW2+ELeq+g12Bp7jJXhRVWkzC2NTJ3L7Owx03Xclk0A80//CWHtnafu1POy3Sxq/lgjwc8JsjqB6/AEd6KWpP82g1vdBM8sFxiHNzObs+qwqOamxcJtm2FKTjRXU6P7Z9IQbT7vpoE64iKu99rjKsuhHK90m59ElYn3uRLlw2NM+2sXQJfDx2tbXAF8FdvES6Gvb159x6BAvQUws5S5q3TQ76Kb7uqnuEHb0FVNQvVgPGBO3gPeYxBuQIf7WBoS8ZWNCeYnnOxnRg3CfDv29kIX7e53OMNijbT/YO6Dwxw97/XY4wmxGtWUX/wF8UK+AbjQAAA==
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
-39
View File
@@ -1,39 +0,0 @@
---
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: grafana
namespace: grafana
labels:
app.kubernetes.io/name: grafana
app.kubernetes.io/instance: grafana
traefik.io/instance: internal
annotations:
cert-manager.io/cluster-issuer: vault-issuer
cert-manager.io/common-name: grafana.k8s.syd1.au.unkin.net
cert-manager.io/private-key-size: "4096"
external-dns.alpha.kubernetes.io/hostname: grafana.k8s.syd1.au.unkin.net
external-dns.alpha.kubernetes.io/target: 198.18.200.4
spec:
gatewayClassName: traefik-internal
listeners:
- name: http
port: 80
protocol: HTTP
hostname: grafana.k8s.syd1.au.unkin.net
allowedRoutes:
namespaces:
from: Same
- name: https
port: 443
protocol: HTTPS
hostname: grafana.k8s.syd1.au.unkin.net
allowedRoutes:
namespaces:
from: Same
tls:
mode: Terminate
certificateRefs:
- group: ""
kind: Secret
name: grafana-tls
-63
View File
@@ -1,63 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: Grafana
metadata:
name: grafana
namespace: grafana
labels:
dashboards: "grafana"
spec:
deployment:
spec:
template:
spec:
containers:
- name: grafana
env:
# DB password + OAuth client secret injected from the
# Vault-synced secrets (GF_ env overrides grafana.ini).
- name: GF_DATABASE_PASSWORD
valueFrom:
secretKeyRef:
name: postgres-credentials
key: password
- name: GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET
valueFrom:
secretKeyRef:
name: oauth-credentials
key: client_secret
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1"
memory: 1Gi
config:
server:
root_url: "https://grafana.k8s.syd1.au.unkin.net"
database:
type: "postgres"
host: "postgres-pooler-rw.grafana.svc.cluster.local:5432"
name: "grafana"
user: "grafana"
ssl_mode: "require"
auth:
disable_login_form: "false"
oauth_auto_login: "false"
"auth.generic_oauth":
enabled: "true"
name: "Authentik"
allow_sign_up: "true"
use_pkce: "true"
client_id: "grafana"
# ak_groups = hierarchical group claim from terraform-authentik (carries
# permission groups inherited via role groups).
scopes: "openid email profile ak_groups"
auth_url: "https://identity.unkin.net/application/o/authorize/"
token_url: "https://identity.unkin.net/application/o/token/"
api_url: "https://identity.unkin.net/application/o/userinfo/"
# Authentik permission groups -> Grafana roles. akP-grafana-admin is granted
# to akR-global-admin members (and direct members) via terraform-authentik.
role_attribute_path: "contains(ak_groups[*], 'akP-grafana-admin') && 'Admin' || 'Viewer'"
role_attribute_strict: "false"
-22
View File
@@ -1,22 +0,0 @@
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDatasource
metadata:
name: victoriametrics
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
# uid matches the previous default datasource so the imported dashboards
# (which hardcode this uid or use the default) resolve without edits.
datasource:
name: "VictoriaMetrics"
type: "prometheus"
uid: "det2y55dac4jkc"
access: "proxy"
url: "http://vmselect-main.observability.svc.cluster.local:8481/select/0/prometheus"
isDefault: true
jsonData:
timeInterval: "15s"
httpMethod: "POST"
-55
View File
@@ -1,55 +0,0 @@
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: grafana-http-redirect
namespace: grafana
labels:
app.kubernetes.io/name: grafana
app.kubernetes.io/instance: grafana
spec:
hostnames:
- grafana.k8s.syd1.au.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: grafana
sectionName: http
rules:
- filters:
- type: RequestRedirect
requestRedirect:
scheme: https
statusCode: 301
matches:
- path:
type: PathPrefix
value: /
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: grafana
namespace: grafana
labels:
app.kubernetes.io/name: grafana
app.kubernetes.io/instance: grafana
spec:
hostnames:
- grafana.k8s.syd1.au.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: grafana
sectionName: https
rules:
- backendRefs:
- group: ""
kind: Service
name: grafana-service
port: 3000
weight: 1
matches:
- path:
type: PathPrefix
value: /
-29
View File
@@ -1,29 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- cnpg_cluster.yaml
- cnpg_backup.yaml
- cnpg_pooler.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
- grafana.yaml
- grafanadatasource.yaml
- gateway.yaml
- httproute.yaml
- dashboards/bind9-exporter-dns.yaml
- dashboards/ceph-cluster.yaml
- dashboards/frr-ospf-route-metrics.yaml
- dashboards/gitea.yaml
- dashboards/haproxy.yaml
- dashboards/media-dashboard.yaml
- dashboards/node-exporter-full.yaml
- dashboards/nzbget.yaml
- dashboards/postgresql-database.yaml
- dashboards/puppet-report.yaml
- dashboards/victorialogs-cluster.yaml
- dashboards/victoriametrics-cluster.yaml
- dashboards/victoriametrics-vmagent.yaml
- dashboards/cloudnativepg.yaml
-7
View File
@@ -1,7 +0,0 @@
---
apiVersion: v1
kind: Namespace
metadata:
labels:
app.kubernetes.io/name: grafana
name: grafana
-18
View File
@@ -1,18 +0,0 @@
---
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultAuth
metadata:
name: default
namespace: grafana
spec:
allowedNamespaces:
- grafana
kubernetes:
audiences:
- vault
role: default
serviceAccount: default
tokenExpirationSeconds: 600
method: kubernetes
mount: k8s/au/syd1
vaultConnectionRef: vso-system/default
-34
View File
@@ -1,34 +0,0 @@
---
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultStaticSecret
metadata:
name: postgres-credentials
namespace: grafana
spec:
destination:
create: true
name: postgres-credentials
overwrite: true
hmacSecretData: true
mount: kv
path: kubernetes/namespace/grafana/default/postgres-credentials
refreshAfter: 5m
type: kv-v2
vaultAuthRef: default
---
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultStaticSecret
metadata:
name: oauth-credentials
namespace: grafana
spec:
destination:
create: true
name: oauth-credentials
overwrite: true
hmacSecretData: true
mount: kv
path: kubernetes/namespace/grafana/default/oauth-credentials
refreshAfter: 5m
type: kv-v2
vaultAuthRef: default
+45
View File
@@ -0,0 +1,45 @@
---
# Ceph RGW (S3) backup target for the jellyfin CNPG cluster, provisioned by the
# in-estate cephrgw-operator: one dedicated bucket + owner user. CNPG reads the
# S3 credential Secret from its own namespace.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-jellyfin-backup
namespace: jellyfin
spec:
displayName: "CNPG backup owner (jellyfin)"
uid: cnpg-jellyfin-backup
maxBuckets: 5
secretName: cnpg-jellyfin-backup-s3
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-jellyfin
namespace: jellyfin
spec:
placementTarget: ec
bucketName: cnpg-jellyfin
ownerRef: cnpg-jellyfin-backup
versioning: false
tags:
app: jellyfin
purpose: cnpg-backup
retainOnDelete: true
---
# Nightly base backup on top of always-on WAL archiving. Scheduled off-peak and
# staggered from the other CNPG clusters (6-field cron, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-jellyfin-nightly
namespace: jellyfin
spec:
schedule: "0 35 3 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: jellyfin-postgres
+119
View File
@@ -0,0 +1,119 @@
---
# Main Jellyfin database. The jellyfin-ha fork's experimental EF Core provider
# moves the entire Jellyfin DB (incl. library items) off SQLite into PostgreSQL,
# which is what makes a shared-nothing multi-replica deployment possible. No
# bootstrap secret is given, so CNPG generates the jellyfin-postgres-app secret
# (username/password/dbname) that the StatefulSet composes its DSN from.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: jellyfin-postgres
namespace: jellyfin
spec:
affinity:
podAntiAffinityType: preferred
backup:
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-jellyfin
endpointURL: https://s3.ceph.unkin.net
endpointCA:
name: vault-ca-cert
key: ca.crt
s3Credentials:
accessKeyId:
name: cnpg-jellyfin-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-jellyfin-backup-s3
key: AWS_SECRET_ACCESS_KEY
serverName: jellyfin
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: jellyfin
encoding: UTF8
localeCType: C
localeCollate: C
owner: jellyfin
enablePDB: true
enableSuperuserAccess: false
failoverDelay: 0
# PG 17 — accepted by the fork's Npgsql/EF Core provider (needs PG14+); the
# provider generates its own migrations on first start.
imageName: ghcr.io/cloudnative-pg/postgresql:17-system-trixie
instances: 3
logLevel: info
maxSyncReplicas: 0
minSyncReplicas: 0
monitoring:
customQueriesConfigMap:
- key: queries
name: cnpg-default-monitoring
disableDefaultQueries: false
enablePodMonitor: false
postgresql:
parameters:
archive_mode: "on"
archive_timeout: 5min
dynamic_shared_memory_type: posix
effective_cache_size: 256MB
full_page_writes: "on"
log_destination: csvlog
log_directory: /controller/log
log_filename: postgres
log_rotation_age: "0"
log_rotation_size: "0"
log_truncate_on_rotation: "false"
logging_collector: "on"
max_connections: "200"
max_parallel_workers: "16"
max_replication_slots: "16"
max_worker_processes: "16"
shared_buffers: 128MB
shared_memory_type: mmap
ssl_max_protocol_version: TLSv1.3
ssl_min_protocol_version: TLSv1.3
wal_keep_size: 256MB
wal_level: logical
wal_log_hints: "on"
wal_receiver_timeout: 5s
wal_sender_timeout: 5s
syncReplicaElectionConstraint:
enabled: false
primaryUpdateMethod: restart
primaryUpdateStrategy: unsupervised
probes:
liveness:
isolationCheck:
connectionTimeout: 1000
enabled: true
requestTimeout: 1000
replicationSlots:
highAvailability:
enabled: true
slotPrefix: _cnpg_
synchronizeReplicas:
enabled: true
updateInterval: 30
resources:
limits:
cpu: "1"
memory: 1Gi
requests:
cpu: 50m
memory: 512Mi
smartShutdownTimeout: 180
startDelay: 3600
stopDelay: 1800
storage:
resizeInUseVolumes: true
size: 10Gi
storageClass: cephrbd-fast-delete
switchoverDelay: 3600
@@ -1,12 +1,15 @@
---
# PgBouncer pooler in front of the jellyfin-postgres cluster. Jellyfin connects
# here (jellyfin-postgres-pooler:5432) rather than the -rw service so EF Core's
# connection churn is absorbed by the pool.
apiVersion: postgresql.cnpg.io/v1
kind: Pooler
metadata:
name: postgres-pooler-rw
namespace: grafana
name: jellyfin-postgres-pooler
namespace: jellyfin
spec:
cluster:
name: postgres
name: jellyfin-postgres
instances: 2
pgbouncer:
parameters:
@@ -17,7 +20,7 @@ spec:
template:
metadata:
labels:
app: pooler-rw
app: jellyfin-pooler
spec:
affinity:
podAntiAffinity:
@@ -27,7 +30,7 @@ spec:
- key: app
operator: In
values:
- pooler-rw
- jellyfin-pooler
topologyKey: kubernetes.io/hostname
containers: []
type: rw
@@ -6,26 +6,26 @@ metadata:
traefik.io/instance: internal
annotations:
cert-manager.io/cluster-issuer: vault-issuer
cert-manager.io/common-name: pdbmux.k8s.syd1.au.unkin.net
cert-manager.io/common-name: jellyfin.k8s.syd1.au.unkin.net
cert-manager.io/private-key-size: "4096"
external-dns.alpha.kubernetes.io/hostname: pdbmux.k8s.syd1.au.unkin.net
external-dns.alpha.kubernetes.io/hostname: jellyfin.k8s.syd1.au.unkin.net
external-dns.alpha.kubernetes.io/target: 198.18.200.4
name: pdbmux
namespace: pdbmux
name: jellyfin
namespace: jellyfin
spec:
gatewayClassName: traefik-internal
listeners:
- allowedRoutes:
namespaces:
from: Same
hostname: pdbmux.k8s.syd1.au.unkin.net
hostname: jellyfin.k8s.syd1.au.unkin.net
name: http
port: 80
protocol: HTTP
- allowedRoutes:
namespaces:
from: Same
hostname: pdbmux.k8s.syd1.au.unkin.net
hostname: jellyfin.k8s.syd1.au.unkin.net
name: https
port: 443
protocol: HTTPS
@@ -33,5 +33,5 @@ spec:
certificateRefs:
- group: ""
kind: Secret
name: pdbmux-tls
name: jellyfin-tls
mode: Terminate
@@ -2,15 +2,15 @@
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: pdbmux-http-redirect
namespace: pdbmux
name: http-redirect
namespace: jellyfin
spec:
hostnames:
- pdbmux.k8s.syd1.au.unkin.net
- jellyfin.k8s.syd1.au.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: pdbmux
name: jellyfin
sectionName: http
rules:
- filters:
@@ -26,22 +26,22 @@ spec:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: pdbmux
namespace: pdbmux
name: jellyfin-route
namespace: jellyfin
spec:
hostnames:
- pdbmux.k8s.syd1.au.unkin.net
- jellyfin.k8s.syd1.au.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: pdbmux
name: jellyfin
sectionName: https
rules:
- backendRefs:
- group: ""
kind: Service
name: pdbmux
port: 80
name: jellyfin
port: 8096
weight: 1
matches:
- path:
@@ -5,12 +5,16 @@ kind: Kustomization
resources:
- namespace.yaml
- cnpg_cluster.yaml
- cnpg_backup.yaml
- cnpg_pooler.yaml
- valkey-deployment.yaml
- valkey-pvc.yaml
- valkey-service.yaml
- vaultauth.yaml
- vaultstaticsecret.yaml
- cnpg_backup.yaml
- pvc-config.yaml
- pvc-transcode.yaml
- pvc-media.yaml
- statefulset.yaml
- pdb.yaml
- service.yaml
- redis-deployment.yaml
- redis-pvc.yaml
- redis-service.yaml
- gateway.yaml
- httproute.yaml
@@ -2,4 +2,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: pdbmux
name: jellyfin
+13
View File
@@ -0,0 +1,13 @@
---
# Keep at least one Jellyfin replica serving through voluntary disruptions
# (node drains, rollouts) so active streams can fail over rather than drop.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: jellyfin
namespace: jellyfin
spec:
minAvailable: 1
selector:
matchLabels:
app: jellyfin
+17
View File
@@ -0,0 +1,17 @@
---
# Jellyfin config: metadata images, plugins, subtitles and config XML. Shared
# ReadWriteMany across replicas (all pods read/write the same library metadata);
# the main library DB now lives in PostgreSQL, not here. Retain — this is state.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-config
namespace: jellyfin
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 20Gi
storageClassName: cephfs-raid5-retain
volumeMode: Filesystem
+17
View File
@@ -0,0 +1,17 @@
---
# Media library, shared read-many across replicas. Retain — this holds the
# actual media and must survive PVC deletion. Empty on first deploy; populating
# it is out of scope for this app.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-media
namespace: jellyfin
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Ti
storageClassName: cephfs-raid6-retain
volumeMode: Filesystem
+18
View File
@@ -0,0 +1,18 @@
---
# Shared transcode scratch. ReadWriteMany is the hard requirement for the HA
# fork: a taking-over pod must read the in-flight HLS segments written by the
# pod it replaces. Scratch data (delete reclaim); raid5 avoids the raid6
# double-parity write penalty on the many small HLS segment writes.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-transcode
namespace: jellyfin
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 100Gi
storageClassName: cephfs-raid5-delete
volumeMode: Filesystem
+66
View File
@@ -0,0 +1,66 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: redis
namespace: jellyfin
spec:
replicas: 1
selector:
matchLabels:
app: redis
strategy:
type: Recreate
template:
metadata:
labels:
app: redis
spec:
containers:
- name: redis
image: redis:7-alpine
imagePullPolicy: IfNotPresent
command:
- redis-server
- --save
- "20"
- "1"
ports:
- containerPort: 6379
name: redis
protocol: TCP
livenessProbe:
exec:
command:
- redis-cli
- ping
failureThreshold: 3
initialDelaySeconds: 30
periodSeconds: 30
successThreshold: 1
timeoutSeconds: 5
readinessProbe:
exec:
command:
- redis-cli
- ping
failureThreshold: 3
initialDelaySeconds: 5
periodSeconds: 10
successThreshold: 1
timeoutSeconds: 5
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 50m
memory: 128Mi
volumeMounts:
- mountPath: /data
name: data
restartPolicy: Always
volumes:
- name: data
persistentVolumeClaim:
claimName: jellyfin-redis-data
@@ -2,8 +2,8 @@
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: netbox-valkey-data
namespace: netbox
name: jellyfin-redis-data
namespace: jellyfin
spec:
accessModes:
- ReadWriteOnce
@@ -2,16 +2,16 @@
apiVersion: v1
kind: Service
metadata:
name: pdbmux
namespace: pdbmux
name: redis
namespace: jellyfin
spec:
internalTrafficPolicy: Cluster
ports:
- name: http
port: 80
- name: redis
port: 6379
protocol: TCP
targetPort: http
targetPort: redis
selector:
app: pdbmux
app: redis
sessionAffinity: None
type: ClusterIP
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: v1
kind: Service
metadata:
name: jellyfin
namespace: jellyfin
spec:
internalTrafficPolicy: Cluster
ports:
- name: http
port: 8096
protocol: TCP
targetPort: http
selector:
app: jellyfin
# Pin each client to one replica to reduce transcode-session churn/takeover.
sessionAffinity: ClientIP
type: ClusterIP
+199
View File
@@ -0,0 +1,199 @@
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: jellyfin
namespace: jellyfin
spec:
# HA: two replicas coordinate transcode session ownership through Redis and
# resume each other's HLS segments off the shared RWX transcode PVC. Stable
# pod names (jellyfin-0/1) are the lease owner identity, hence StatefulSet.
replicas: 2
serviceName: jellyfin
podManagementPolicy: Parallel
updateStrategy:
type: RollingUpdate
selector:
matchLabels:
app: jellyfin
template:
metadata:
labels:
app: jellyfin
spec:
securityContext:
# Group-write the shared RWX volumes and grant the render/video groups so
# the runAsUser 1000 process can open the Intel DRI render node injected
# by the device plugin.
fsGroup: 1000
supplementalGroups:
- 44
- 105
- 109
seccompProfile:
type: RuntimeDefault
affinity:
# Spread the two replicas across nodes for node-level HA. Soft so a
# single-GPU-node cluster still schedules both (i915 has 4 shared slots).
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: jellyfin
topologyKey: kubernetes.io/hostname
initContainers:
# Select the fork's experimental PostgreSQL provider by writing
# database.xml before Jellyfin starts. Runs as root to chown into the
# shared config volume; mirrors the fork Helm chart's inject-db-config.
- name: inject-db-config
image: busybox:1.37.0
command:
- sh
- -c
- |
mkdir -p /config/config
chown 1000:1000 /config/config
chmod 775 /config/config
cat > /config/config/database.xml << 'DBEOF'
<?xml version="1.0" encoding="utf-8"?>
<DatabaseConfigurationOptions>
<DatabaseType>Jellyfin-PostgreSQL</DatabaseType>
<LockingBehavior>NoLock</LockingBehavior>
</DatabaseConfigurationOptions>
DBEOF
chown 1000:1000 /config/config/database.xml
chmod 664 /config/config/database.xml
resources:
requests:
cpu: 10m
memory: 32Mi
limits:
cpu: 100m
memory: 64Mi
volumeMounts:
- name: config
mountPath: /config
containers:
- name: jellyfin
image: git.unkin.net/unkin/jellyfin-ha:v0.1.0
imagePullPolicy: IfNotPresent
ports:
- name: http
containerPort: 8096
protocol: TCP
env:
# Pod identity for the Redis transcode lease owner. The fork reads
# JELLYFIN_INSTANCE_ID (falling back to MachineName); the stable
# StatefulSet pod name gives each replica a unique lease identity so
# takeover can target a dead replica. JELLYFIN_HA_POD_NAME is set for
# parity with the fork Helm chart (nothing currently reads it).
- name: JELLYFIN_INSTANCE_ID
valueFrom:
fieldRef:
fieldPath: metadata.name
- name: JELLYFIN_HA_POD_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
# Multiple replicas must not each answer UDP auto-discovery.
- name: JELLYFIN_Network__AutoDiscovery
value: "false"
# Config dir must differ from the data root (Jellyfin sanity check).
- name: JELLYFIN_CONFIG_DIR
value: /config/config
# Distributed transcode session store (jellyfin-ha additions).
- name: Jellyfin__TranscodeStore__RedisConnectionString
value: "redis:6379,abortConnect=false"
- name: Jellyfin__TranscodeStore__LeaseDurationSeconds
value: "30"
# PostgreSQL main DB via the CNPG-generated app secret, routed through
# the PgBouncer pooler. Composed with $(VAR) expansion so the password
# is never rendered into the manifest; CNPG passwords are URL-safe.
- name: PGUSER
valueFrom:
secretKeyRef:
name: jellyfin-postgres-app
key: username
- name: PGPASSWORD
valueFrom:
secretKeyRef:
name: jellyfin-postgres-app
key: password
- name: PGDB
valueFrom:
secretKeyRef:
name: jellyfin-postgres-app
key: dbname
- name: POSTGRES_CONNECTION_STRING
value: "postgresql://$(PGUSER):$(PGPASSWORD)@jellyfin-postgres-pooler:5432/$(PGDB)"
- name: DATABASE_URL
value: "postgresql://$(PGUSER):$(PGPASSWORD)@jellyfin-postgres-pooler:5432/$(PGDB)"
livenessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
resources:
requests:
cpu: "1"
memory: 1Gi
gpu.intel.com/i915: "1"
limits:
cpu: "4"
memory: 6Gi
# Intel iGPU (QSV/VA-API) slot. Requesting it pins the pod to a
# GPU-labelled node and injects /dev/dri/renderD* automatically, so
# no /dev/dri hostPath or privileged container is needed. Enable
# QSV/VA-API once in the Jellyfin admin UI; it persists to /config.
gpu.intel.com/i915: "1"
securityContext:
runAsUser: 1000
runAsGroup: 1000
volumeMounts:
- name: config
mountPath: /config
- name: transcode
# Fork's real transcode temp path. RWX so a surviving pod reads the
# in-flight .ts/.m3u8 segments of the pod it takes over. A per-pod
# volume here silently breaks HA takeover.
mountPath: /config/transcodes
- name: cache
mountPath: /cache
- name: media
mountPath: /media
readOnly: true
volumes:
- name: config
persistentVolumeClaim:
claimName: jellyfin-config
- name: transcode
persistentVolumeClaim:
claimName: jellyfin-transcode
- name: media
persistentVolumeClaim:
claimName: jellyfin-media
volumeClaimTemplates:
# Per-pod scratch cache — RWO, disposable, one PVC per replica.
- metadata:
name: cache
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 30Gi
storageClassName: cephrbd-fast-delete
volumeMode: Filesystem
-1
View File
@@ -14,7 +14,6 @@ resources:
- gateway.yaml
- httproute.yaml
- tlsroute.yaml
- vpa.yaml
configMapGenerator:
- name: kanidm-config
-13
View File
@@ -1,13 +0,0 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: kanidm-vpa
namespace: kanidm
spec:
targetRef:
apiVersion: apps/v1
kind: StatefulSet
name: kanidm
updatePolicy:
updateMode: "Off"
-55
View File
@@ -1,55 +0,0 @@
---
# Ceph RGW (S3) backup target for the litellm CNPG cluster, provisioned by the
# in-estate cephrgw-operator. One dedicated bucket + owner user per cluster:
# cephrgw CRs are namespace-scoped and CNPG reads its S3 credential Secret from
# its own namespace, so backups are per-database rather than one shared bucket.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-litellm-backup
namespace: litellm
spec:
displayName: "CNPG backup owner (litellm)"
# RGW users are global; keep the uid namespace-qualified so it never collides.
uid: cnpg-litellm-backup
maxBuckets: 5
# Operator writes AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (+ RGW_UID,
# S3_ENDPOINT) into this Secret; the Cluster's barmanObjectStore consumes it.
secretName: cnpg-litellm-backup-s3
# Keep the RGW user (and thus the keys) if this CR is ever deleted, so an
# in-flight restore can still reach the archive.
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-litellm
namespace: litellm
spec:
bucketName: cnpg-litellm
# The owner user has full control of its own bucket (read + write), which is
# all the backup/restore identity needs — no extra BucketAccess grant.
ownerRef: cnpg-litellm-backup
versioning: false
tags:
app: litellm
purpose: cnpg-backup
# Never drop the backups if the CR is removed; retire buckets by hand.
retainOnDelete: true
---
# Nightly base backup. Continuous WAL archiving is always-on via the Cluster's
# spec.backup.barmanObjectStore; this schedules the periodic full backup that
# WAL is layered on top of. Schedules are staggered across clusters so the 8
# base backups do not hit RGW at once (CNPG cron is 6-field, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-litellm-nightly
namespace: litellm
spec:
schedule: "0 20 1 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: litellm-postgres
+1 -30
View File
@@ -7,35 +7,6 @@ metadata:
spec:
affinity:
podAntiAffinityType: preferred
backup:
# 30-day retention (DEFAULT — adjust per cluster if needed). Enforced by CNPG
# against the object store on each successful base backup.
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-litellm
endpointURL: https://s3.ceph.unkin.net
# radosgw serves a Vault-PKI cert; trust the internal CA (reflected into
# every namespace as the vault-ca-cert Secret).
endpointCA:
name: vault-ca-cert
key: ca.crt
# Keys minted by the ObjectStoreUser in cnpg_backup.yaml; never hardcoded.
s3Credentials:
accessKeyId:
name: cnpg-litellm-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-litellm-backup-s3
key: AWS_SECRET_ACCESS_KEY
# Path prefix within the bucket; keep stable across restores (see docs).
serverName: litellm
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: litellm
@@ -108,7 +79,7 @@ spec:
cpu: "1"
memory: 1Gi
requests:
cpu: 50m
cpu: 250m
memory: 512Mi
smartShutdownTimeout: 180
startDelay: 3600
-8
View File
@@ -30,14 +30,6 @@ spec:
- containerPort: 4000
name: http
protocol: TCP
env:
# Authentik OIDC client secret (SSO); non-secret SSO config is in the
# litellm-env ConfigMap.
- name: GENERIC_CLIENT_SECRET
valueFrom:
secretKeyRef:
name: oauth-credentials
key: client_secret
envFrom:
- secretRef:
name: litellm-credentials

Some files were not shown because too many files have changed in this diff Show More