Deploy jellyfin-ha as a true-HA StatefulSet under a new media project #237

Merged
benvin merged 10 commits from benvin/jellyfin-ha into main 2026-08-15 13:39:00 +10:00

10 Commits

Author SHA1 Message Date
unkin-agent f5a721cd0b jellyfin: back media PVCs with shared static CephFS PVs
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
2026-08-15 13:03:49 +10:00
unkin-agent 066dc41b55 jellyfin: scope k8up backup to config PVC only
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
The jellyfin-config Schedule had no PVC selector and the k8up chart runs
skipWithoutAnnotation=false, so restic swept in every PVC in the namespace
(media, transcode scratch, redis, per-pod cache, CNPG data) — RWO volumes
also fail to mount while in use. k8up 4.10.0's Schedule CRD has no positive
PVC selector, so exclude every non-config PVC by annotation.

- Annotate media-tv, media-movies, transcode and redis-data PVCs with
  k8up.io/backup: "false".
- Annotate the per-pod cache volumeClaimTemplate the same way.
- Propagate the annotation to the CNPG data PVCs via inheritedMetadata
  (postgres has its own barmanObjectStore backup).
- Leave jellyfin-config unannotated so it remains the only backup target.
2026-08-15 12:34:25 +10:00
unkin-agent 57ac95bdec jellyfin: back up the config PVC with a k8up Schedule
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Protect the jellyfin-config PVC (library metadata, plugins, config XML)
with daily restic backups to the new Ceph RGW config-backup bucket.

- Add a k8up.io Schedule: daily backup (02:00), weekly prune (Sun 03:00,
  keep 14 daily/8 weekly/12 monthly) and weekly check (Sun 04:00)
- Source S3 creds from the cephrgw BucketAccess Secret and the restic
  repo password from Vault via a VaultStaticSecret (kv path
  kubernetes/namespace/jellyfin/default/k8up-restic)
- Add the namespace VaultAuth (default role/SA) VSO needs to sync it
- Mount the reflected vault-ca-cert into the restic pods so restic trusts
  the internal unkin.net CA on s3.ceph.unkin.net
- Wire the new files into the jellyfin kustomization
2026-08-15 12:08:50 +10:00
unkin-agent 6b30958a24 jellyfin: add cephrgw config-backup bucket for k8up
Give the jellyfin backup user a second bucket so k8up can store restic
backups of the jellyfin-config PVC alongside the existing CNPG barman
bucket (one user, two buckets).

- Add a jellyfin-config-backup Bucket owned by the existing
  cnpg-jellyfin-backup ObjectStoreUser
- Add a read-write BucketAccess emitting jellyfin-config-backup-s3 with
  the S3 creds k8up consumes
- Wire the new file into the jellyfin kustomization
2026-08-15 12:06:30 +10:00
unkin-agent 1f63cc370d Merge remote-tracking branch 'origin/main' into benvin/jellyfin-ha 2026-08-15 12:01:40 +10:00
unkin-agent 6769b5291e jellyfin: split media PVC into tv/movies on cephfs-raid5-delete
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Separate tv and movies onto their own volumes so sonarr/radarr can each
mount and manage their library individually later, and move media storage
off cephfs-raid6-retain onto cephfs-raid5-delete.

- Remove single jellyfin-media PVC (RWX, cephfs-raid6-retain)
- Add jellyfin-media-tv and jellyfin-media-movies PVCs (RWX, cephfs-raid5-delete, 500Gi each, expandable)
- Mount tv read-only at /media/tv and movies read-only at /media/movies in the statefulset
- Update kustomization resources to reference the two new PVCs
2026-08-15 11:55:55 +10:00
Ben Vincent 648ff16940 jellyfin-ha: pull image from artifactapi docker-internal registry
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
The build repo's release pipeline pushes the runtime image to the
artifactapi local docker registry, not Gitea. Point the StatefulSet at
artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.1.0 so
the deploy pulls the image the build actually produces. Pull is anonymous
in-cluster (no imagePullSecret), matching every other docker-internal
workload in the estate (encapi, pdbmux, bind-operator, cephrgw, age-api).
2026-08-11 07:36:23 +10:00
Ben Vincent ceb2114467 Enable Intel iGPU hardware transcode by default for jellyfin-ha
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
The pod already requests gpu.intel.com/i915 and joins the render/video
groups, but Jellyfin transcodes in software until hardware acceleration
is turned on in its encoding config, which the fork does not template.

Seed /config/config/encoding.xml from the init container so VAAPI on the
injected Intel render node (/dev/dri/renderD128) is active on first boot:
HardwareAccelerationType vaapi, EnableHardwareEncoding, h264/hevc decode,
tonemapping left off. Rename the init container to inject-config and write
each seed file only when absent, so admin changes persisted to the shared
RWX /config are never clobbered on restart.
2026-08-11 07:23:38 +10:00
Ben Vincent f89e1c8260 Rework jellyfin-ha into a true-HA StatefulSet deployment
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Turn the single-replica jellyfin-ha app into a proper high-availability
deployment so the fork's Redis-coordinated distributed transcoding and
PostgreSQL main database can actually be exercised.

- Replace the Deployment with a 2-replica StatefulSet for stable pod
  identity; set JELLYFIN_INSTANCE_ID from metadata.name (the fork's Redis
  transcode-lease owner id), add soft podAntiAffinity and a PDB
  minAvailable 1.
- Move the main Jellyfin DB to PostgreSQL via a CloudNativePG trio
  (3-instance Cluster, PgBouncer Pooler, Ceph RGW barman backups) mirroring
  the litellm pattern; an init container writes database.xml selecting the
  fork's Jellyfin-PostgreSQL provider and the DSN is composed from the
  CNPG-generated app secret pointed at the pooler.
- Share /config on an RWX cephfs PVC across replicas; keep /cache per-pod
  via a volumeClaimTemplate.
- Fix the transcode mount to the fork's real path /config/transcodes on the
  RWX PVC (raid5) so a surviving pod can resume the segments of the pod it
  takes over.
- Add Intel iGPU hardware transcoding via the gpu.intel.com/i915 device
  plugin resource plus render/video supplemental groups.
- Switch the Service to sessionAffinity ClientIP to reduce transcode churn.
- Disable UDP auto-discovery. Library scans still run on every replica; a
  single-scanner leader election is a planned follow-up.
2026-08-10 23:36:53 +10:00
Ben Vin ed75a6d1e1 Add jellyfin (HA fork) app under a new media project
Deploys the jellyfin-ha fork (git.unkin.net/unkin/jellyfin-ha) to au-syd1
via ArgoCD, under a dedicated media AppProject/ApplicationSet rather than
extending platform.

- New media AppProject + media-apps ApplicationSet (watches
  apps/overlays/*/jellyfin); registered in the argocd kustomizations
- apps/base/jellyfin: namespace, deployment (single replica to start),
  service, in-namespace Redis (transcode session store), gateway + httproute
  at jellyfin.k8s.syd1.au.unkin.net
- Storage: RWO config (cephrbd), RWX transcode scratch and RWX media
  library (cephfs) per the HA fork's pod-takeover requirement
- au-syd1 overlay
2026-08-10 23:36:53 +10:00