Compare commits

..

3 Commits

Author SHA1 Message Date
Ben Vincent 54c25be828 Reduce media PR to jellyfin-only in its own namespace
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Why:
Jellyfin ships and gets validated first, ahead of the rest of the media stack.
Scoping this PR to jellyfin alone keeps the initial rollout small and lets the
HA fork prove out against the real library before the download and manager apps
follow.

How:
- Drop sonarr, radarr, prowlarr, bazarr, nzbget, and jellyseerr and their shared
  media-apps foundation from this PR; they land in later PRs.
- Move jellyfin into its own jellyfin namespace and fold the namespace and the
  static mediafs PV plus its RWX claim into the jellyfin base.
- Keep the static CephFS PV bound to the in-use mediafs library with
  reclaimPolicy Retain and staticVolume true so nothing can reclaim it, mounted
  into jellyfin by the movies and tvseries subPaths; keep redis, the fresh RWX
  transcode scratch, the intel iGPU nodeSelector and i915 request, gateway, and
  httproute.
- Scope the media AppProject and ApplicationSet to the single jellyfin
  namespace and app, extensible as the remaining apps are added.
2026-08-09 21:08:12 +10:00
Ben Vincent a52a419dfd Point media apps at the real mediafs library via a static CephFS PV
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Why:
The media apps must serve and manage the actual media library, not empty
volumes. That library already exists on the puppet-managed CephFS filesystem
mediafs (mounted by the VM/incus instances at /shared/media) and is in active
use, so the k8s apps must mount it in place rather than provision fresh storage.

How:
- Replace the two fresh movies/tvseries PVCs with one static CephFS
  PersistentVolume bound to mediafs and a single RWX media-library claim the
  whole stack shares.
- Set the PV reclaim policy to Retain and mark it staticVolume so ceph-csi only
  mounts the pre-existing storage and can never provision or reclaim it;
  deleting the PVC or PV cannot destroy the underlying library.
- Reuse the live csi-cephfs cluster parameters (clusterID cephfs_csi_ssd_ec_4_1
  for mon discovery, csi-cephfs/csi-cephfs-secret node-stage secret) with
  fsName mediafs and rootPath / (the mediafs root that maps to /shared/media).
- Mount the library into each app by subPath so the tree matches the VM
  layout: sonarr /mnt/tvseries (tvseries), radarr /mnt/movies (movies),
  jellyfin and nzbget both subtrees; prowlarr keeps no library mount. The
  jellyfin transcode PVC stays a fresh scratch volume.
- Whitelist PersistentVolume in the media AppProject so the cluster-scoped PV
  can sync.
2026-08-09 13:39:40 +10:00
Ben Vincent e03aeca101 Add media-apps stack to ArgoCD
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Why:
The media stack (jellyfin plus the sonarr/radarr/prowlarr/bazarr/nzbget/
jellyseerr apps) runs in the media-apps namespace but is deployed out-of-band
by terraform-k8s rather than GitOps. Bringing it under ArgoCD makes the stack
declarative, self-healing, and consistent with every other cluster workload,
and prepares terraform-k8s to drop the media-apps config.

How:
- Add a media AppProject scoped to the media-apps namespace and a media-apps
  ApplicationSet that renders one Application per app plus a shared foundation.
- Add a shared media-apps foundation (namespace, media-apps-vault-reader
  ServiceAccount, default VaultAuth on k8s/au/syd1, and the RWX movies/tvseries
  library PVCs) that the whole stack mounts.
- Add per-app kustomize base and au-syd1 overlay for jellyfin and the six *arr
  apps, using plain resource names (jellyfin, sonarr, ...) with fresh PVCs.
- Deploy jellyfin from the jellyfin-ha fork (Redis transcode store, RWX
  transcode scratch) wired to the shared movies/tvseries library PVCs, keeping
  the intel iGPU nodeSelector and gpu.intel.com/i915 request.
- Source API keys and nzbget credentials through VSO VaultStaticSecrets from
  kv/service/media-apps/<app>; expose each app via a traefik-internal Gateway
  and HTTPRoute at <app>.k8s.syd1.au.unkin.net.
- Register the media project and applicationset in the argocd bootstrap
  kustomizations.
2026-08-09 13:25:25 +10:00
24 changed files with 236 additions and 510 deletions
+1 -1
View File
@@ -36,7 +36,7 @@ spec:
mountPath: /combined-certs
containers:
- name: api
image: git.unkin.net/unkin/artifactapi:v3.8.0
image: git.unkin.net/unkin/artifactapi:v3.7.7
imagePullPolicy: IfNotPresent
ports:
- containerPort: 8000
+1 -1
View File
@@ -22,7 +22,7 @@ spec:
automountServiceAccountToken: true
containers:
- name: ui
image: git.unkin.net/unkin/artifactapi-ui:v3.8.0
image: git.unkin.net/unkin/artifactapi-ui:v3.7.7
imagePullPolicy: IfNotPresent
ports:
- containerPort: 80
-45
View File
@@ -1,45 +0,0 @@
---
# Ceph RGW (S3) backup target for the jellyfin CNPG cluster, provisioned by the
# in-estate cephrgw-operator: one dedicated bucket + owner user. CNPG reads the
# S3 credential Secret from its own namespace.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: cnpg-jellyfin-backup
namespace: jellyfin
spec:
displayName: "CNPG backup owner (jellyfin)"
uid: cnpg-jellyfin-backup
maxBuckets: 5
secretName: cnpg-jellyfin-backup-s3
retainOnDelete: true
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: cnpg-jellyfin
namespace: jellyfin
spec:
placementTarget: ec
bucketName: cnpg-jellyfin
ownerRef: cnpg-jellyfin-backup
versioning: false
tags:
app: jellyfin
purpose: cnpg-backup
retainOnDelete: true
---
# Nightly base backup on top of always-on WAL archiving. Scheduled off-peak and
# staggered from the other CNPG clusters (6-field cron, seconds first).
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: cnpg-jellyfin-nightly
namespace: jellyfin
spec:
schedule: "0 35 3 * * *"
immediate: false
backupOwnerReference: self
method: barmanObjectStore
cluster:
name: jellyfin-postgres
-119
View File
@@ -1,119 +0,0 @@
---
# Main Jellyfin database. The jellyfin-ha fork's experimental EF Core provider
# moves the entire Jellyfin DB (incl. library items) off SQLite into PostgreSQL,
# which is what makes a shared-nothing multi-replica deployment possible. No
# bootstrap secret is given, so CNPG generates the jellyfin-postgres-app secret
# (username/password/dbname) that the StatefulSet composes its DSN from.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: jellyfin-postgres
namespace: jellyfin
spec:
affinity:
podAntiAffinityType: preferred
backup:
retentionPolicy: 30d
barmanObjectStore:
# Dedicated per-cluster Ceph RGW bucket (cephrgw-operator provisions it).
destinationPath: s3://cnpg-jellyfin
endpointURL: https://s3.ceph.unkin.net
endpointCA:
name: vault-ca-cert
key: ca.crt
s3Credentials:
accessKeyId:
name: cnpg-jellyfin-backup-s3
key: AWS_ACCESS_KEY_ID
secretAccessKey:
name: cnpg-jellyfin-backup-s3
key: AWS_SECRET_ACCESS_KEY
serverName: jellyfin
data:
compression: bzip2
jobs: 2
wal:
compression: zstd
maxParallel: 2
bootstrap:
initdb:
database: jellyfin
encoding: UTF8
localeCType: C
localeCollate: C
owner: jellyfin
enablePDB: true
enableSuperuserAccess: false
failoverDelay: 0
# PG 17 — accepted by the fork's Npgsql/EF Core provider (needs PG14+); the
# provider generates its own migrations on first start.
imageName: ghcr.io/cloudnative-pg/postgresql:17-system-trixie
instances: 3
logLevel: info
maxSyncReplicas: 0
minSyncReplicas: 0
monitoring:
customQueriesConfigMap:
- key: queries
name: cnpg-default-monitoring
disableDefaultQueries: false
enablePodMonitor: false
postgresql:
parameters:
archive_mode: "on"
archive_timeout: 5min
dynamic_shared_memory_type: posix
effective_cache_size: 256MB
full_page_writes: "on"
log_destination: csvlog
log_directory: /controller/log
log_filename: postgres
log_rotation_age: "0"
log_rotation_size: "0"
log_truncate_on_rotation: "false"
logging_collector: "on"
max_connections: "200"
max_parallel_workers: "16"
max_replication_slots: "16"
max_worker_processes: "16"
shared_buffers: 128MB
shared_memory_type: mmap
ssl_max_protocol_version: TLSv1.3
ssl_min_protocol_version: TLSv1.3
wal_keep_size: 256MB
wal_level: logical
wal_log_hints: "on"
wal_receiver_timeout: 5s
wal_sender_timeout: 5s
syncReplicaElectionConstraint:
enabled: false
primaryUpdateMethod: restart
primaryUpdateStrategy: unsupervised
probes:
liveness:
isolationCheck:
connectionTimeout: 1000
enabled: true
requestTimeout: 1000
replicationSlots:
highAvailability:
enabled: true
slotPrefix: _cnpg_
synchronizeReplicas:
enabled: true
updateInterval: 30
resources:
limits:
cpu: "1"
memory: 1Gi
requests:
cpu: 50m
memory: 512Mi
smartShutdownTimeout: 180
startDelay: 3600
stopDelay: 1800
storage:
resizeInUseVolumes: true
size: 10Gi
storageClass: cephrbd-fast-delete
switchoverDelay: 3600
-36
View File
@@ -1,36 +0,0 @@
---
# PgBouncer pooler in front of the jellyfin-postgres cluster. Jellyfin connects
# here (jellyfin-postgres-pooler:5432) rather than the -rw service so EF Core's
# connection churn is absorbed by the pool.
apiVersion: postgresql.cnpg.io/v1
kind: Pooler
metadata:
name: jellyfin-postgres-pooler
namespace: jellyfin
spec:
cluster:
name: jellyfin-postgres
instances: 2
pgbouncer:
parameters:
default_pool_size: "50"
max_client_conn: "200"
paused: false
poolMode: session
template:
metadata:
labels:
app: jellyfin-pooler
spec:
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app
operator: In
values:
- jellyfin-pooler
topologyKey: kubernetes.io/hostname
containers: []
type: rw
+97
View File
@@ -0,0 +1,97 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: jellyfin
namespace: jellyfin
spec:
# Single-replica for now. The jellyfin-ha fork adds the Redis transcode store
# and RWX transcode scratch that make scaling to true HA a follow-up.
replicas: 1
strategy:
# Config PVC is RWO; Recreate avoids two pods contending for it.
type: Recreate
selector:
matchLabels:
app: jellyfin
template:
metadata:
labels:
app: jellyfin
spec:
securityContext:
fsGroup: 1000
nodeSelector:
feature.node.kubernetes.io/pci-0300_8086.present: "true"
containers:
- name: jellyfin
image: git.unkin.net/unkin/jellyfin-ha:v0.1.0
imagePullPolicy: IfNotPresent
ports:
- name: http
containerPort: 8096
protocol: TCP
env:
- name: TZ
value: Australia/Sydney
- name: PUID
value: "1000"
- name: PGID
value: "1000"
- name: JELLYFIN_PublishedServerUrl
value: https://jellyfin.k8s.syd1.au.unkin.net
# Distributed transcode session store (jellyfin-ha additions).
- name: Jellyfin__TranscodeStore__RedisConnectionString
value: "jellyfin-redis:6379,abortConnect=false"
- name: Jellyfin__TranscodeStore__LeaseDurationSeconds
value: "30"
resources:
requests:
cpu: 100m
memory: 1Gi
limits:
cpu: "4"
memory: 8Gi
gpu.intel.com/i915: "1"
livenessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
volumeMounts:
- name: config
mountPath: /config
- name: cache
mountPath: /cache
- name: transcode
mountPath: /transcode
- name: media-library
mountPath: /mnt/movies
subPath: movies
- name: media-library
mountPath: /mnt/tvseries
subPath: tvseries
volumes:
- name: config
persistentVolumeClaim:
claimName: jellyfin-config
- name: cache
persistentVolumeClaim:
claimName: jellyfin-cache
- name: transcode
persistentVolumeClaim:
claimName: jellyfin-transcode
- name: media-library
persistentVolumeClaim:
claimName: media-library
+1 -1
View File
@@ -26,7 +26,7 @@ spec:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: jellyfin-route
name: jellyfin
namespace: jellyfin
spec:
hostnames:
+7 -9
View File
@@ -4,17 +4,15 @@ kind: Kustomization
resources:
- namespace.yaml
- cnpg_cluster.yaml
- cnpg_pooler.yaml
- cnpg_backup.yaml
- pvc-config.yaml
- pvc-transcode.yaml
- pvc-media.yaml
- statefulset.yaml
- pdb.yaml
- pv_media-library.yaml
- pvc_media-library.yaml
- pvc_config.yaml
- pvc_cache.yaml
- pvc_transcode.yaml
- deployment.yaml
- service.yaml
- redis-deployment.yaml
- redis-pvc.yaml
- redis-service.yaml
- redis-pvc.yaml
- gateway.yaml
- httproute.yaml
-13
View File
@@ -1,13 +0,0 @@
---
# Keep at least one Jellyfin replica serving through voluntary disruptions
# (node drains, rollouts) so active streams can fail over rather than drop.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: jellyfin
namespace: jellyfin
spec:
minAvailable: 1
selector:
matchLabels:
app: jellyfin
+41
View File
@@ -0,0 +1,41 @@
---
# Static CephFS PersistentVolume bound to the pre-existing, ACTIVELY-USED
# puppet media library (ceph filesystem "mediafs", mounted by the VM/incus
# instances at /shared/media). ceph-csi only mounts this volume; staticVolume
# tells it the storage pre-exists and it must never provision or delete it.
#
# reclaimPolicy MUST stay Retain: deleting this PV or its PVC must NEVER be able
# to reclaim or destroy the underlying CephFS data that the VM instances use.
apiVersion: v1
kind: PersistentVolume
metadata:
name: jellyfin-media-library
spec:
accessModes:
- ReadWriteMany
capacity:
storage: 10Ti
# Load-bearing safety control. Do not change to Delete.
persistentVolumeReclaimPolicy: Retain
storageClassName: ""
volumeMode: Filesystem
# Pre-bind to the media-library claim so nothing else can grab this PV.
claimRef:
apiVersion: v1
kind: PersistentVolumeClaim
name: media-library
namespace: jellyfin
csi:
driver: cephfs.csi.ceph.com
volumeHandle: jellyfin-media-library-static
nodeStageSecretRef:
name: csi-cephfs-secret
namespace: csi-cephfs
volumeAttributes:
# clusterID maps (in the csi-cephfs ceph-csi-config) to the mon set that
# also serves mediafs; for a static volume only the mon lookup is used.
clusterID: cephfs_csi_ssd_ec_4_1
fsName: mediafs
staticVolume: "true"
# Filesystem-internal root of the library (mediafs root == /shared/media).
rootPath: /
-17
View File
@@ -1,17 +0,0 @@
---
# Jellyfin config: metadata images, plugins, subtitles and config XML. Shared
# ReadWriteMany across replicas (all pods read/write the same library metadata);
# the main library DB now lives in PostgreSQL, not here. Retain — this is state.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-config
namespace: jellyfin
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 20Gi
storageClassName: cephfs-raid5-retain
volumeMode: Filesystem
-17
View File
@@ -1,17 +0,0 @@
---
# Media library, shared read-many across replicas. Retain — this holds the
# actual media and must survive PVC deletion. Empty on first deploy; populating
# it is out of scope for this app.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-media
namespace: jellyfin
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Ti
storageClassName: cephfs-raid6-retain
volumeMode: Filesystem
+15
View File
@@ -0,0 +1,15 @@
---
# Local transcode/image cache. Scratch, delete reclaim policy.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-cache
namespace: jellyfin
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 200Gi
storageClassName: cephrbd-fast-delete
volumeMode: Filesystem
+15
View File
@@ -0,0 +1,15 @@
---
# Jellyfin config + SQLite library database. Single-writer, block storage.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: jellyfin-config
namespace: jellyfin
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 10Gi
storageClassName: cephrbd-fast-retain
volumeMode: Filesystem
+19
View File
@@ -0,0 +1,19 @@
---
# Claim bound to the static mediafs PV. RWX so every app in the stack shares the
# one library. storageClassName "" + volumeName pin it to the static PV (no
# dynamic provisioning). Deleting this claim cannot reclaim the data (PV is
# Retain).
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: media-library
namespace: jellyfin
spec:
accessModes:
- ReadWriteMany
storageClassName: ""
volumeName: jellyfin-media-library
resources:
requests:
storage: 10Ti
volumeMode: Filesystem
@@ -1,8 +1,7 @@
---
# Shared transcode scratch. ReadWriteMany is the hard requirement for the HA
# fork: a taking-over pod must read the in-flight HLS segments written by the
# pod it replaces. Scratch data (delete reclaim); raid5 avoids the raid6
# double-parity write penalty on the many small HLS segment writes.
# pod it replaces. Scratch data, so delete reclaim policy.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
+15 -17
View File
@@ -2,20 +2,21 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: redis
name: jellyfin-redis
namespace: jellyfin
spec:
replicas: 1
selector:
matchLabels:
app: redis
strategy:
type: Recreate
selector:
matchLabels:
app: jellyfin-redis
template:
metadata:
labels:
app: redis
app: jellyfin-redis
spec:
restartPolicy: Always
containers:
- name: redis
image: redis:7-alpine
@@ -26,40 +27,37 @@ spec:
- "20"
- "1"
ports:
- containerPort: 6379
name: redis
- name: redis
containerPort: 6379
protocol: TCP
livenessProbe:
exec:
command:
- redis-cli
- ping
failureThreshold: 3
initialDelaySeconds: 30
periodSeconds: 30
successThreshold: 1
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
exec:
command:
- redis-cli
- ping
failureThreshold: 3
initialDelaySeconds: 5
periodSeconds: 10
successThreshold: 1
timeoutSeconds: 5
failureThreshold: 3
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 50m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
volumeMounts:
- mountPath: /data
name: data
restartPolicy: Always
- name: data
mountPath: /data
volumes:
- name: data
persistentVolumeClaim:
+6 -6
View File
@@ -2,16 +2,16 @@
apiVersion: v1
kind: Service
metadata:
name: redis
name: jellyfin-redis
namespace: jellyfin
spec:
type: ClusterIP
internalTrafficPolicy: Cluster
sessionAffinity: None
selector:
app: jellyfin-redis
ports:
- name: redis
port: 6379
protocol: TCP
targetPort: redis
selector:
app: redis
sessionAffinity: None
type: ClusterIP
protocol: TCP
+5 -6
View File
@@ -5,14 +5,13 @@ metadata:
name: jellyfin
namespace: jellyfin
spec:
type: ClusterIP
internalTrafficPolicy: Cluster
sessionAffinity: None
selector:
app: jellyfin
ports:
- name: http
port: 8096
protocol: TCP
targetPort: http
selector:
app: jellyfin
# Pin each client to one replica to reduce transcode-session churn/takeover.
sessionAffinity: ClientIP
type: ClusterIP
protocol: TCP
-199
View File
@@ -1,199 +0,0 @@
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: jellyfin
namespace: jellyfin
spec:
# HA: two replicas coordinate transcode session ownership through Redis and
# resume each other's HLS segments off the shared RWX transcode PVC. Stable
# pod names (jellyfin-0/1) are the lease owner identity, hence StatefulSet.
replicas: 2
serviceName: jellyfin
podManagementPolicy: Parallel
updateStrategy:
type: RollingUpdate
selector:
matchLabels:
app: jellyfin
template:
metadata:
labels:
app: jellyfin
spec:
securityContext:
# Group-write the shared RWX volumes and grant the render/video groups so
# the runAsUser 1000 process can open the Intel DRI render node injected
# by the device plugin.
fsGroup: 1000
supplementalGroups:
- 44
- 105
- 109
seccompProfile:
type: RuntimeDefault
affinity:
# Spread the two replicas across nodes for node-level HA. Soft so a
# single-GPU-node cluster still schedules both (i915 has 4 shared slots).
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: jellyfin
topologyKey: kubernetes.io/hostname
initContainers:
# Select the fork's experimental PostgreSQL provider by writing
# database.xml before Jellyfin starts. Runs as root to chown into the
# shared config volume; mirrors the fork Helm chart's inject-db-config.
- name: inject-db-config
image: busybox:1.37.0
command:
- sh
- -c
- |
mkdir -p /config/config
chown 1000:1000 /config/config
chmod 775 /config/config
cat > /config/config/database.xml << 'DBEOF'
<?xml version="1.0" encoding="utf-8"?>
<DatabaseConfigurationOptions>
<DatabaseType>Jellyfin-PostgreSQL</DatabaseType>
<LockingBehavior>NoLock</LockingBehavior>
</DatabaseConfigurationOptions>
DBEOF
chown 1000:1000 /config/config/database.xml
chmod 664 /config/config/database.xml
resources:
requests:
cpu: 10m
memory: 32Mi
limits:
cpu: 100m
memory: 64Mi
volumeMounts:
- name: config
mountPath: /config
containers:
- name: jellyfin
image: git.unkin.net/unkin/jellyfin-ha:v0.1.0
imagePullPolicy: IfNotPresent
ports:
- name: http
containerPort: 8096
protocol: TCP
env:
# Pod identity for the Redis transcode lease owner. The fork reads
# JELLYFIN_INSTANCE_ID (falling back to MachineName); the stable
# StatefulSet pod name gives each replica a unique lease identity so
# takeover can target a dead replica. JELLYFIN_HA_POD_NAME is set for
# parity with the fork Helm chart (nothing currently reads it).
- name: JELLYFIN_INSTANCE_ID
valueFrom:
fieldRef:
fieldPath: metadata.name
- name: JELLYFIN_HA_POD_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
# Multiple replicas must not each answer UDP auto-discovery.
- name: JELLYFIN_Network__AutoDiscovery
value: "false"
# Config dir must differ from the data root (Jellyfin sanity check).
- name: JELLYFIN_CONFIG_DIR
value: /config/config
# Distributed transcode session store (jellyfin-ha additions).
- name: Jellyfin__TranscodeStore__RedisConnectionString
value: "redis:6379,abortConnect=false"
- name: Jellyfin__TranscodeStore__LeaseDurationSeconds
value: "30"
# PostgreSQL main DB via the CNPG-generated app secret, routed through
# the PgBouncer pooler. Composed with $(VAR) expansion so the password
# is never rendered into the manifest; CNPG passwords are URL-safe.
- name: PGUSER
valueFrom:
secretKeyRef:
name: jellyfin-postgres-app
key: username
- name: PGPASSWORD
valueFrom:
secretKeyRef:
name: jellyfin-postgres-app
key: password
- name: PGDB
valueFrom:
secretKeyRef:
name: jellyfin-postgres-app
key: dbname
- name: POSTGRES_CONNECTION_STRING
value: "postgresql://$(PGUSER):$(PGPASSWORD)@jellyfin-postgres-pooler:5432/$(PGDB)"
- name: DATABASE_URL
value: "postgresql://$(PGUSER):$(PGPASSWORD)@jellyfin-postgres-pooler:5432/$(PGDB)"
livenessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
resources:
requests:
cpu: "1"
memory: 1Gi
gpu.intel.com/i915: "1"
limits:
cpu: "4"
memory: 6Gi
# Intel iGPU (QSV/VA-API) slot. Requesting it pins the pod to a
# GPU-labelled node and injects /dev/dri/renderD* automatically, so
# no /dev/dri hostPath or privileged container is needed. Enable
# QSV/VA-API once in the Jellyfin admin UI; it persists to /config.
gpu.intel.com/i915: "1"
securityContext:
runAsUser: 1000
runAsGroup: 1000
volumeMounts:
- name: config
mountPath: /config
- name: transcode
# Fork's real transcode temp path. RWX so a surviving pod reads the
# in-flight .ts/.m3u8 segments of the pod it takes over. A per-pod
# volume here silently breaks HA takeover.
mountPath: /config/transcodes
- name: cache
mountPath: /cache
- name: media
mountPath: /media
readOnly: true
volumes:
- name: config
persistentVolumeClaim:
claimName: jellyfin-config
- name: transcode
persistentVolumeClaim:
claimName: jellyfin-transcode
- name: media
persistentVolumeClaim:
claimName: jellyfin-media
volumeClaimTemplates:
# Per-pod scratch cache — RWO, disposable, one PVC per replica.
- metadata:
name: cache
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 30Gi
storageClassName: cephrbd-fast-delete
volumeMode: Filesystem
@@ -4,8 +4,6 @@ metadata:
annotations:
configmap.reloader.stakater.com/auto: "true"
secret.reloader.stakater.com/reload: "vault-ca-cert"
# Replace clears the stale, API-server-defaulted spec.strategy.rollingUpdate that SSA cannot drop, which otherwise makes Recreate invalid.
argocd.argoproj.io/sync-options: Replace=true
labels:
app.kubernetes.io/component: puppetserver
app.kubernetes.io/instance: puppetserver
+8 -7
View File
@@ -6,11 +6,12 @@ metadata:
namespace: argocd
spec:
generators:
- git:
repoURL: https://git.unkin.net/unkin/argocd-apps
revision: HEAD
directories:
- path: apps/overlays/*/jellyfin
- git:
repoURL: https://git.unkin.net/unkin/argocd-apps
revision: HEAD
directories:
# jellyfin only for now; downloads/managers apps join in later PRs.
- path: apps/overlays/*/jellyfin
template:
metadata:
name: 'media-{{path[3]}}'
@@ -22,10 +23,10 @@ spec:
path: '{{path}}'
destination:
server: https://kubernetes.default.svc
namespace: '{{path[3]}}'
namespace: jellyfin
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- ServerSideApply=true
- ServerSideApply=true
+4 -2
View File
@@ -5,15 +5,17 @@ metadata:
name: media
namespace: argocd
spec:
description: Media services
description: Media services (jellyfin; downloads/managers namespaces to follow)
sourceRepos:
- https://git.unkin.net/unkin/argocd-apps
destinations:
- namespace: 'jellyfin'
- namespace: jellyfin
server: https://kubernetes.default.svc
clusterResourceWhitelist:
- group: ''
kind: Namespace
- group: ''
kind: PersistentVolume
namespaceResourceWhitelist:
- group: '*'
kind: '*'
@@ -6,16 +6,6 @@ metadata:
namespace: argocd
data:
kustomize.buildOptions: "--enable-helm"
# Kubernetes defaults apiVersion/kind onto every StatefulSet
# volumeClaimTemplates entry, but neither the raw manifests nor the helm
# charts emit them, so live StatefulSets carry TypeMeta that git lacks.
# volumeClaimTemplates are immutable on an existing StatefulSet, so ArgoCD
# can never reconcile the removal and the resource stays perpetually
# OutOfSync. Ignore the defaulted TypeMeta fleet-wide.
resource.customizations.ignoreDifferences.apps_StatefulSet: |
jqPathExpressions:
- '.spec.volumeClaimTemplates[]?.apiVersion'
- '.spec.volumeClaimTemplates[]?.kind'
# External URL ArgoCD serves on (TLS terminated at the traefik-internal gateway).
url: https://argocd.k8s.syd1.au.unkin.net
# OIDC login via Authentik. The client secret is seeded in Vault out of band