6c8c0dd9e4
## Why Deploy the jellyfin-ha fork as a genuine high-availability service rather than a single replica, so its two headline capabilities can actually be exercised: the Redis-coordinated distributed transcoding (a surviving pod resumes the HLS segments of a pod that dies mid-stream) and the experimental PostgreSQL main database (which is what lets more than one replica share the same library). It lands in its own `jellyfin` namespace under a new `media` ArgoCD project. ## How **Workload — StatefulSet, 2 replicas.** The Deployment becomes a StatefulSet so each pod has a stable name. The fork's Redis transcode lease keys ownership on `JELLYFIN_INSTANCE_ID`, which is set from the downward-API pod name, giving each replica a unique, stable lease identity. Soft `podAntiAffinity` spreads the two pods across nodes and a `PodDisruptionBudget` keeps `minAvailable: 1` through drains and rollouts. **Main database — CloudNativePG.** A CNPG trio in-namespace mirrors the litellm pattern: a 3-instance `Cluster`, a PgBouncer `Pooler`, and Ceph RGW (barman) S3 backups to a dedicated `cnpg-jellyfin` bucket owned by a `cephrgw` `ObjectStoreUser`. An init container writes `/config/config/database.xml` selecting the fork's `Jellyfin-PostgreSQL` provider, and the connection string is composed from the CNPG-generated `jellyfin-postgres-app` secret (username / password / dbname) pointed at the pooler service — the password is never rendered into the manifest. Library-item metadata therefore moves off SQLite; metadata images, plugins, subtitles and config XML stay on `/config`. **Storage.** - `/config` is now a shared `ReadWriteMany` cephfs PVC (raid5, retain) so every replica reads/writes the same metadata and config. - `/config/transcodes` — the fork's real transcode temp path — is a shared RWX PVC (raid5, delete). This is the load-bearing fix: takeover reads the dead pod's in-flight `.ts`/`.m3u8` segments off shared storage, so per-pod scratch would silently break it. - `/cache` is per-pod via a `volumeClaimTemplate` (RWO). - The media library stays a fresh, empty RWX PVC mounted read-only; populating it is out of scope. **Hardware transcoding.** The container requests the `gpu.intel.com/i915` Intel device-plugin resource (which pins the pod to a GPU-labelled node and injects the DRI render node — no `/dev/dri` hostPath or privileged container) plus the render/video supplemental groups. VA-API hardware acceleration is now on by default: the `inject-config` init container seeds `/config/config/encoding.xml` with `HardwareAccelerationType` `vaapi`, `EnableHardwareEncoding`, the injected render node (`/dev/dri/renderD128`) and h264/hevc hardware decode, so transcodes use the iGPU on first boot with no manual admin-UI step. Both seed files (`database.xml`, `encoding.xml`) are written only when absent, so later admin changes persisted to the shared RWX `/config` are never clobbered on restart. **Networking.** The Gateway/HTTPRoute (traefik-internal, `jellyfin.k8s.syd1.au.unkin.net`) is unchanged; the Service gains `sessionAffinity: ClientIP` to keep a client pinned to one replica and reduce transcode-session churn. **Redis.** The in-namespace single-replica Redis stays as the transcode lease store. ## Follow-up UDP auto-discovery is disabled, but scheduled library scans still run on every replica (redundant scans). Single-scanner leader election is a planned follow-up pending a fork feature, tracked separately. --------- Co-authored-by: unkin-agent <unkin-agent@unkin.net> Co-authored-by: Ben Vincent <neotheo@gmail.com> Co-authored-by: Ben Vin <neotheo@gmail.com> Reviewed-on: #237 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net>
246 lines
9.8 KiB
YAML
246 lines
9.8 KiB
YAML
---
|
|
apiVersion: apps/v1
|
|
kind: StatefulSet
|
|
metadata:
|
|
name: jellyfin
|
|
namespace: jellyfin
|
|
spec:
|
|
# HA: two replicas coordinate transcode session ownership through Redis and
|
|
# resume each other's HLS segments off the shared RWX transcode PVC. Stable
|
|
# pod names (jellyfin-0/1) are the lease owner identity, hence StatefulSet.
|
|
replicas: 2
|
|
serviceName: jellyfin
|
|
podManagementPolicy: Parallel
|
|
updateStrategy:
|
|
type: RollingUpdate
|
|
selector:
|
|
matchLabels:
|
|
app: jellyfin
|
|
template:
|
|
metadata:
|
|
labels:
|
|
app: jellyfin
|
|
spec:
|
|
securityContext:
|
|
# Group-write the shared RWX volumes and grant the render/video groups so
|
|
# the runAsUser 1000 process can open the Intel DRI render node injected
|
|
# by the device plugin.
|
|
fsGroup: 1000
|
|
supplementalGroups:
|
|
- 44
|
|
- 105
|
|
- 109
|
|
seccompProfile:
|
|
type: RuntimeDefault
|
|
affinity:
|
|
# Spread the two replicas across nodes for node-level HA. Soft so a
|
|
# single-GPU-node cluster still schedules both (i915 has 4 shared slots).
|
|
podAntiAffinity:
|
|
preferredDuringSchedulingIgnoredDuringExecution:
|
|
- weight: 100
|
|
podAffinityTerm:
|
|
labelSelector:
|
|
matchLabels:
|
|
app: jellyfin
|
|
topologyKey: kubernetes.io/hostname
|
|
initContainers:
|
|
# Seed the fork's PostgreSQL provider (database.xml) and Intel iGPU
|
|
# hardware transcode settings (encoding.xml) before Jellyfin starts.
|
|
# Runs as root to chown into the shared config volume; mirrors the fork
|
|
# Helm chart's inject-db-config. Each file is written only when absent so
|
|
# admin changes persisted to the shared RWX /config survive pod restarts.
|
|
- name: inject-config
|
|
image: busybox:1.37.0
|
|
command:
|
|
- sh
|
|
- -c
|
|
- |
|
|
mkdir -p /config/config
|
|
chown 1000:1000 /config/config
|
|
chmod 775 /config/config
|
|
if [ ! -e /config/config/database.xml ]; then
|
|
cat > /config/config/database.xml << 'DBEOF'
|
|
<?xml version="1.0" encoding="utf-8"?>
|
|
<DatabaseConfigurationOptions>
|
|
<DatabaseType>Jellyfin-PostgreSQL</DatabaseType>
|
|
<LockingBehavior>NoLock</LockingBehavior>
|
|
</DatabaseConfigurationOptions>
|
|
DBEOF
|
|
chown 1000:1000 /config/config/database.xml
|
|
chmod 664 /config/config/database.xml
|
|
fi
|
|
# VAAPI on the Intel render node the device plugin injects
|
|
# (/dev/dri/renderD128 — ffmpeg's default DRM node, reachable via
|
|
# the render/video supplementalGroups). Without this the attached
|
|
# iGPU is idle and every transcode runs in software. Omitted
|
|
# elements fall back to the fork's EncodingOptions defaults.
|
|
if [ ! -e /config/config/encoding.xml ]; then
|
|
cat > /config/config/encoding.xml << 'ENCEOF'
|
|
<?xml version="1.0" encoding="utf-8"?>
|
|
<EncodingOptions xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xsd="http://www.w3.org/2001/XMLSchema">
|
|
<EncodingThreadCount>-1</EncodingThreadCount>
|
|
<HardwareAccelerationType>vaapi</HardwareAccelerationType>
|
|
<VaapiDevice>/dev/dri/renderD128</VaapiDevice>
|
|
<EnableHardwareEncoding>true</EnableHardwareEncoding>
|
|
<AllowHevcEncoding>true</AllowHevcEncoding>
|
|
<AllowAv1Encoding>false</AllowAv1Encoding>
|
|
<EnableIntelLowPowerH264HwEncoder>false</EnableIntelLowPowerH264HwEncoder>
|
|
<EnableIntelLowPowerHevcHwEncoder>false</EnableIntelLowPowerHevcHwEncoder>
|
|
<EnableTonemapping>false</EnableTonemapping>
|
|
<EnableVppTonemapping>false</EnableVppTonemapping>
|
|
<HardwareDecodingCodecs>
|
|
<string>h264</string>
|
|
<string>hevc</string>
|
|
<string>vc1</string>
|
|
<string>vp9</string>
|
|
</HardwareDecodingCodecs>
|
|
</EncodingOptions>
|
|
ENCEOF
|
|
chown 1000:1000 /config/config/encoding.xml
|
|
chmod 664 /config/config/encoding.xml
|
|
fi
|
|
resources:
|
|
requests:
|
|
cpu: 10m
|
|
memory: 32Mi
|
|
limits:
|
|
cpu: 100m
|
|
memory: 64Mi
|
|
volumeMounts:
|
|
- name: config
|
|
mountPath: /config
|
|
containers:
|
|
- name: jellyfin
|
|
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.1.0
|
|
imagePullPolicy: IfNotPresent
|
|
ports:
|
|
- name: http
|
|
containerPort: 8096
|
|
protocol: TCP
|
|
env:
|
|
# Pod identity for the Redis transcode lease owner. The fork reads
|
|
# JELLYFIN_INSTANCE_ID (falling back to MachineName); the stable
|
|
# StatefulSet pod name gives each replica a unique lease identity so
|
|
# takeover can target a dead replica. JELLYFIN_HA_POD_NAME is set for
|
|
# parity with the fork Helm chart (nothing currently reads it).
|
|
- name: JELLYFIN_INSTANCE_ID
|
|
valueFrom:
|
|
fieldRef:
|
|
fieldPath: metadata.name
|
|
- name: JELLYFIN_HA_POD_NAME
|
|
valueFrom:
|
|
fieldRef:
|
|
fieldPath: metadata.name
|
|
# Multiple replicas must not each answer UDP auto-discovery.
|
|
- name: JELLYFIN_Network__AutoDiscovery
|
|
value: "false"
|
|
# Config dir must differ from the data root (Jellyfin sanity check).
|
|
- name: JELLYFIN_CONFIG_DIR
|
|
value: /config/config
|
|
# Distributed transcode session store (jellyfin-ha additions).
|
|
- name: Jellyfin__TranscodeStore__RedisConnectionString
|
|
value: "redis:6379,abortConnect=false"
|
|
- name: Jellyfin__TranscodeStore__LeaseDurationSeconds
|
|
value: "30"
|
|
# PostgreSQL main DB via the CNPG-generated app secret, routed through
|
|
# the PgBouncer pooler. Composed with $(VAR) expansion so the password
|
|
# is never rendered into the manifest; CNPG passwords are URL-safe.
|
|
- name: PGUSER
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: jellyfin-postgres-app
|
|
key: username
|
|
- name: PGPASSWORD
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: jellyfin-postgres-app
|
|
key: password
|
|
- name: PGDB
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: jellyfin-postgres-app
|
|
key: dbname
|
|
- name: POSTGRES_CONNECTION_STRING
|
|
value: "postgresql://$(PGUSER):$(PGPASSWORD)@jellyfin-postgres-pooler:5432/$(PGDB)"
|
|
- name: DATABASE_URL
|
|
value: "postgresql://$(PGUSER):$(PGPASSWORD)@jellyfin-postgres-pooler:5432/$(PGDB)"
|
|
livenessProbe:
|
|
httpGet:
|
|
path: /health
|
|
port: http
|
|
initialDelaySeconds: 30
|
|
periodSeconds: 30
|
|
timeoutSeconds: 5
|
|
failureThreshold: 3
|
|
readinessProbe:
|
|
httpGet:
|
|
path: /health
|
|
port: http
|
|
initialDelaySeconds: 10
|
|
periodSeconds: 10
|
|
timeoutSeconds: 5
|
|
failureThreshold: 3
|
|
resources:
|
|
requests:
|
|
cpu: "1"
|
|
memory: 1Gi
|
|
gpu.intel.com/i915: "1"
|
|
limits:
|
|
cpu: "4"
|
|
memory: 6Gi
|
|
# Intel iGPU (QSV/VA-API) slot. Requesting it pins the pod to a
|
|
# GPU-labelled node and injects /dev/dri/renderD* automatically, so
|
|
# no /dev/dri hostPath or privileged container is needed. VA-API is
|
|
# pre-enabled via the seeded encoding.xml (see inject-config), so
|
|
# transcodes use the iGPU on first boot with no manual UI step.
|
|
gpu.intel.com/i915: "1"
|
|
securityContext:
|
|
runAsUser: 1000
|
|
runAsGroup: 1000
|
|
volumeMounts:
|
|
- name: config
|
|
mountPath: /config
|
|
- name: transcode
|
|
# Fork's real transcode temp path. RWX so a surviving pod reads the
|
|
# in-flight .ts/.m3u8 segments of the pod it takes over. A per-pod
|
|
# volume here silently breaks HA takeover.
|
|
mountPath: /config/transcodes
|
|
- name: cache
|
|
mountPath: /cache
|
|
- name: media-tv
|
|
mountPath: /media/tv
|
|
readOnly: true
|
|
- name: media-movies
|
|
mountPath: /media/movies
|
|
readOnly: true
|
|
volumes:
|
|
- name: config
|
|
persistentVolumeClaim:
|
|
claimName: jellyfin-config
|
|
- name: transcode
|
|
persistentVolumeClaim:
|
|
claimName: jellyfin-transcode
|
|
- name: media-tv
|
|
persistentVolumeClaim:
|
|
claimName: jellyfin-media-tv
|
|
- name: media-movies
|
|
persistentVolumeClaim:
|
|
claimName: jellyfin-media-movies
|
|
volumeClaimTemplates:
|
|
# Per-pod scratch cache — RWO, disposable, one PVC per replica.
|
|
- metadata:
|
|
name: cache
|
|
annotations:
|
|
# Exclude the per-pod cache PVCs from the jellyfin-config k8up Schedule
|
|
# (skipWithoutAnnotation is false cluster-wide). Cache is disposable and
|
|
# RWO — it would also fail to mount into the backup pod while in use.
|
|
k8up.io/backup: "false"
|
|
spec:
|
|
accessModes:
|
|
- ReadWriteOnce
|
|
resources:
|
|
requests:
|
|
storage: 30Gi
|
|
storageClassName: cephrbd-fast-delete
|
|
volumeMode: Filesystem
|