Commit Graph

463 Commits

Author SHA1 Message Date
unkin-agent cdaab736b5 Bump jellyfin-ha to v0.4.0 (#483)
The deployed v0.3.3 build returns 500 from /Shows/NextUp on PostgreSQL, breaking the home screen, and lets replicas diverge: library-visibility and shared-config changes never propagate, user data (resume, played state, favourites, ratings) is overwritten between pods, and eight scheduled tasks run on every replica instead of only the scan leader. v0.4.0 carries the fixes.

- Pin cheeztv and fafflix to jellyfin-ha:v0.4.0.

No config change needed: cross-pod invalidation reuses the transcode-store Redis connection string both apps already set.

Reviewed-on: #483
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-22 22:57:06 +10:00
unkin-agent b31517e6d9 Merge pull request #481 from benvin/jellyfin-sso-valkey-state
Roll jellyfin-ha to v0.3.3 and drop Service session affinity

Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 14:02:48 +10:00
unkin-agent 5a74b2cec6 Roll bind-operator to v0.2.7 (journal-aware zone seeding) (#480)
Deploy bind-operator v0.2.7. The operator seeded a fresh skeleton zone file at serial 1 over zones whose BIND journal was still on disk at a higher serial; BIND rejected the inconsistent pair (`addzone failed: out of range`) and, with a PVC per replica, the stale journal outlived restarts while every reconcile rewrote the skeleton, so it never converged. That SERVFAILed roughly 1 in 3 authoritative answers for k8s.syd1.au.unkin.net and resolvers cached the failures.

- Bumps the operator image to v0.2.7
- Bumps the CRD install pin to the v0.2.7 tag, which changes the CRDs

Expect one rolling restart of the operator Deployment as the new image lands.

Reviewed-on: #480
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 00:53:47 +10:00
unkin-agent f14bcc4d2a Add woodpecker ServiceAccount for jellyfin-plugin-sso CI (#479)
The new `unkin/jellyfin-plugin-sso` fork is getting a Woodpecker pipeline, and its build step will set `serviceAccountName: jellyfin-plugin-sso`. Without the SA declared here the pipeline pods fail to schedule.

- add a bare ServiceAccount `jellyfin-plugin-sso` in the `woodpecker` namespace
- register it in the woodpecker base kustomization

The step only builds .NET code, so no Vault kube-auth role or RBAC is needed.

Reviewed-on: #479
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 17:32:31 +10:00
unkin-agent b01e4c3241 Rewrite slash-less Authentik token endpoint to the canonical path (#478)
Authentik advertises the token endpoint with a trailing slash, but some OIDC clients (the ArgoCD iOS app) POST to /application/o/token without one; Django's APPEND_SLASH will not redirect a POST, so the token exchange gets 405 and login fails.

- Add an exact-match rule on /application/o/token to the authentik and authentik-internal HTTPRoutes.
- Rewrite it to /application/o/token/ with a URLRewrite ReplaceFullPath filter, preserving the method and the authentik-server backend.
- Leave the catch-all PathPrefix rule untouched; exact matches outrank it in Gateway API precedence.

Reviewed-on: #478
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 16:10:30 +10:00
unkin-agent bbd5bdaa95 Enable PKCE for ArgoCD OIDC login (#477)
The Authentik client for ArgoCD is now public (the iOS app can't hold
a secret), so Authentik no longer enforces client_secret on token
exchange. PKCE replaces that as the protection against
authorization-code interception.

- Add `enablePKCEAuthentication: true` to the `oidc.config` block in
  `argocd-cm-patch.yaml`
- Note why PKCE is needed now that the client is public

Reviewed-on: #477
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 16:10:05 +10:00
unkin-agent 4762cf9e03 Add VMPodScrape for authentik-server metrics (#476)
Authentik server pods expose django_prometheus metrics on :9300, but only ldap-outpost and redis-exporter are scraped in this namespace. Add the missing per-app scrape.

- add apps/base/authentik/server-vmpodscrape.yaml selecting app.kubernetes.io/name=authentik, component=server on the metrics port
- wire it into apps/base/authentik/kustomization.yaml

Reviewed-on: #476
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 15:34:26 +10:00
unkin-agent c83a886e74 Enable pg_stat_statements on the authentik postgres cluster (#475)
The cluster preloads no statement-statistics library, so there is no per-query cost attribution in postgres and slow query paths have to be inferred from application-side metrics instead of read straight out of the database.

- preload `pg_stat_statements`
- set `pg_stat_statements.max` and `.track`, which is what makes CNPG manage the extension and create it in every database

Requires a postgres restart. Stacked on `benvin/authentik-cnpg-resources`.

Reviewed-on: #475
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:48:45 +10:00
unkin-agent 9a7200636c Raise authentik postgres CPU, memory and buffer sizing (#474)
The 500m CPU limit is a 50ms CFS quota per 100ms period, so the postgres pods are throttled on every burst even at ~0.01 cores average and each query pays that latency. 128MB of `shared_buffers` and a 256MB cache estimate also under-serve the planner on the joins authentik issues on its hot read paths.

- raise resources to requests `500m`/`1Gi`, limits `2`/`2Gi`
- raise `shared_buffers` to 512MB and `effective_cache_size` to 1536MB
- hold the post-incident memory headroom multiple over `shared_buffers`

Rolling restart with switchover. Stacked on `benvin/authentik-hot-standby-feedback`.

Reviewed-on: #474
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:35:16 +10:00
unkin-agent 9535bad9bc Enable hot_standby_feedback on the authentik postgres cluster (#473)
Authentik serves multi-second API reads from the CNPG hot standbys. Those reads outlive `max_standby_streaming_delay`, so recovery cancels them with `canceling statement due to conflict with recovery`, which authentik surfaces as HTTP 500 — enough to break a terraform apply mid-run.

- set `hot_standby_feedback` on so replicas report their oldest xmin to the primary and long reads stop being cancelled
- SIGHUP reload only, no restart or switchover
- retained-dead-tuple cost is negligible on a ~155MB database

Reviewed-on: #473
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:34:33 +10:00
unkin-agent 34dd70435e Auto-reload cheeztv and fafflix on plugin ConfigMap change (#471)
Edits to the cheeztv/fafflix plugin ConfigMaps only reach the pods via the inject-plugin-config initContainer, so a config change sat inert until someone manually rolled the StatefulSet. Reloader is deployed cluster-wide with autoReloadAll disabled, so each workload has to opt in.

- annotate both StatefulSets with configmap.reloader.stakater.com/auto: "true"

Reviewed-on: #471
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 12:25:08 +10:00
unkin-agent 6cc752336e Point jellyfin SSO at public Authentik hostname (#470)
The internal-CA identity.k8s.syd1.au.unkin.net host has no CA bundle mounted in the jellyfin pods, so the OIDC discovery fetch fails TLS handshake (PartialChain). Authentik's discovery response is host-relative, so the browser-facing hostname must be used, not the internal one.

- Change OidEndpoint to identity.unkin.net in fafflix plugin config
- Change OidEndpoint to identity.unkin.net in cheeztv plugin config

Reviewed-on: #470
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 12:09:31 +10:00
unkin-agent 47a2ab9152 Pin jellyfin-ha image to v0.3.2 (#468)
v0.3.0 and v0.3.1 crash-looped on Postgres migration/reader bugs and were reverted. v0.3.2 fixes both and was validated end to end against production-baseline Postgres and valkey: full migration chain completes, all previously-500 endpoints return 200, RedisTranscodeSessionStore and scan-leader gating confirmed active.

- Bump jellyfin-ha image tag v0.2.0 -> v0.3.2 in cheeztv and fafflix statefulsets

Depends on a pre-sync duplicate-username check and fresh pg_dump of both databases.

Reviewed-on: #468
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 23:08:44 +10:00
unkin-agent 4748df497a puppet: install toml into the puppetserver gem path (#467)
Catalog compiles fail with `LoadError: no such file to load -- toml`: server-side functions run in the puppetserver JRuby, whose gem path is separate from the agent CRuby path this hook installs into. puppet-prod's `profiles::puppet::gems` covers both; the hook only did the agent half.

- Install toml via `puppetserver gem`, mirroring the `puppetserver_gem` resource in puppet-prod
- Note in a comment that under `set -e` a failed install takes down an already-serving compiler

Reviewed-on: #467
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 23:08:31 +10:00
unkin-agent ba14f85e51 Pin pdbmux to v0.4.0 (#466)
v0.3.0 still serves facts from cache and collapses non-4xx upstream rejections into a 502, so clients see stale facts and lose the real status.

- Pin the pdbmux image to v0.4.0

Reviewed-on: #466
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 20:36:49 +10:00
unkin-agent 8a00ddb82c Revert jellyfin-ha to v0.2.0 (#465)
v0.3.1 crash-loops both jellyfin StatefulSets deterministically on the
RatingLevels migration (concurrent Npgsql command in progress), failing
before any schema change commits. OrderedReady updates leave ordinal-1
stuck, stranding cheeztv and fafflix single-replica with no HA.

- revert cheeztv jellyfin-ha image v0.3.1 -> v0.2.0
- revert fafflix jellyfin-ha image v0.3.1 -> v0.2.0

Unblocks the stalled StatefulSet rollout.

Reviewed-on: #465
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 21:38:08 +10:00
unkin-agent d4aed39f6a pdbmux: bump image to v0.3.0 (#464)
The deployed pin sits on v0.2.0, so pdbmux still answers malformed queries with `502 all backends failed` and resolves per-certname routes by configured backend order rather than by which backend actually owns the node.

Bump the pdbmux image pin to v0.3.0:

- Replay a unanimous upstream rejection (PuppetDB's real 400 + parse message) instead of a 502.
- Resolve per-certname routes to the node's owning backend by report freshness.
- Match backend addresses case-insensitively when redacting.

Reviewed-on: #464
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 20:45:14 +10:00
unkin-agent d6a1279efe jellyfin: bump fafflix+cheeztv to v0.3.1 (#463)
v0.3.0 (#461) crash-looped existing databases on a broken Postgres
migration path; v0.3.1 restores the migration baseline, hardens guid/FK
handling, and fixes encoding.xml parsing.

- bump cheeztv jellyfin-ha image v0.2.0 -> v0.3.1
- bump fafflix jellyfin-ha image v0.2.0 -> v0.3.1

Requires manual pre-merge database verification before merge.

Reviewed-on: #463
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 20:32:41 +10:00
unkin-agent 55af4b2f16 allow catalog-diff to compile catalogs on the puppet compilers (#462)
catalog-diff compiles a host's catalog in two environments and diffs them to validate puppet-prod changes before merge, which means compiling catalogs on behalf of other nodes via POST /puppet/v4/catalog. The compilers run the image default auth.conf, where that endpoint is denied.

- add a compiler auth.conf allowing catalog-diff.main.unkin.net to POST /puppet/v4/catalog
- add a pre-default entrypoint script seeding it into conf.d, failing hard if the source is absent
- mount both onto the compiler deployment via configMapGenerator

Reviewed-on: #462
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 16:31:08 +10:00
unkin-agent 783a3db0fd revert jellyfin-ha to v0.2.0 (#461)
v0.3.0 fails an EF-Core migration on boot (NormalizedUsername column
missing), crash-looping ordinal-1 pods so the StatefulSet rolling
update stalls and cheeztv/fafflix stay single-replica. Unblocks the
stuck rollout.

- revert cheeztv statefulset image to jellyfin-ha:v0.2.0
- revert fafflix statefulset image to jellyfin-ha:v0.2.0

Reviewed-on: #461
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 15:22:20 +10:00
unkin-agent 84f09f89ff bump jellyfin-ha to v0.3.0 (#460)
jellyfin-ha v0.3.0 is the first build tracking Jellyfin 12.0 (.NET 10 runtime, jellyfin-web 12.0, LDAP plugin 24).

- bump cheeztv jellyfin-ha image tag to v0.3.0
- bump fafflix jellyfin-ha image tag to v0.3.0

Reviewed-on: #460
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 14:23:39 +10:00
unkin-agent 5b07157eeb woodpecker: raise agent workspace PVC to 20Gi (#459)
Woodpecker's k8s backend provisions a per-pipeline workspace PVC sized by
WOODPECKER_BACKEND_K8S_VOLUME_SIZE. At 10G, large builds (e.g. .NET clone +
build output) leave too little free space for tests that hard-require free
disk headroom, failing purely on disk exhaustion.

- raise WOODPECKER_BACKEND_K8S_VOLUME_SIZE from 10G to 20Gi

Reviewed-on: #459
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 12:47:07 +10:00
unkin-agent 7aec9a9021 artifactapi: restore combine-certs + PROVIDER_CA_FILES on oauth2-proxy (#458)
**Fix-forward companion to the #457 rollback. This is NOT the current outage fix — see below.**

## The actual outage

The UI is 503 because the Authentik application slug `artifactapi` **does not exist**. OIDC discovery 404s, so oauth2-proxy exits at startup, the Service has no ready endpoints, and Traefik answers `no available server`.

```
identity.unkin.net              /application/o/artifactapi/…  404
identity.k8s.syd1.au.unkin.net  /application/o/artifactapi/…  404
identity.unkin.net              /application/o/repospawner/…  200
identity.unkin.net              /application/o/argocd/…       200
```

Root cause is upstream in **terraform-authentik**: `ci/woodpecker/push/apply` on main HEAD `4e16401` **failed**. That apply has to succeed before any argocd-apps change can help. **This PR does not fix that.**

## What this PR does fix

#456 dropped the `combine-certs` initContainer and `OAUTH2_PROXY_PROVIDER_CA_FILES`, reasoning that `identity.unkin.net` serves a publicly trusted Let's Encrypt cert and so needs no internal CA. That holds for the browser redirect but not for oauth2-proxy's own back-channel discovery/token calls.

artifactapi is the **only one of six** oauth2-proxies in the estate without it:

| app | issuer host | `PROVIDER_CA_FILES` |
|---|---|---|
| arrproxy | identity.unkin.net | yes |
| logviewer | identity.unkin.net | yes |
| mediamark | identity.unkin.net | yes |
| repospawner | identity.unkin.net | yes |
| watchstate | identity.k8s… | yes |
| **artifactapi** | identity.unkin.net | **no** |

repospawner uses the **same public `identity.unkin.net` issuer** and still needs the internal bundle, which falsifies the removal reasoning. The existing comment on that initContainer states it plainly: *"The Authentik issuer is served behind the internal unkin.net CA."*

## Changes

- Add the `combine-certs` initContainer — byte-identical to repospawner's.
- Mount the combined bundle and set `OAUTH2_PROXY_PROVIDER_CA_FILES`.
- Reload the Deployment when `vault-ca-cert` rotates.

`vault-ca-cert` already exists in the `artifactapi` namespace (`api-deployment.yaml` uses it). `kustomize build apps/base/artifactapi` succeeds.

## Risk

Trust-only and strictly additive — it appends the internal CA to the system roots. Harmless if the back channel turns out to reach a publicly trusted endpoint after all. Expected to remove the *next* blocker, surfacing as x509, once the terraform-authentik apply lands.

## Sequencing

1. Fix and re-run terraform-authentik `push/apply` so the `artifactapi` application exists.
2. Merge this.
3. Confirm `/ui/` returns 200, then close #457 unmerged.

Only merge #457 instead if the UI must come back before step 1 can be done.

Reviewed-on: #458
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-07 23:05:39 +10:00
unkin-agent c98d88c197 Put the artifactapi web UI behind Authentik oauth2-proxy (#456)
The artifactapi web UI is open to anyone who can reach the host. Front it with Authentik SSO gated on akP-artifactapi-admin, while leaving the package-manager surfaces (/api/v1, /api/v2, /v2 docker registry, /terraform, /.well-known) untouched — dnf, containerd mirrors, buildah, terraform and CI publish steps cannot do a browser flow.

- Add the oauth2-proxy ConfigMap, Deployment, Service and VMPodScrape.
- Add the oauth-credentials VaultStaticSecret.
- Point the api-route /ui rule at oauth2-proxy and add a /oauth2 rule; the catch-all / rule still goes straight to the api Service on both listeners.

Requires terraform-authentik #34 applied and the Vault kv seed first.

Reviewed-on: #456
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-07 21:01:17 +10:00
unkin-agent 520da449c9 pdbmux: bump image to v0.2.0 (#455)
## Why
v0.2.0 ships the reports/events union, aggregate summing across backends, config-file support, and the removal of the primary/prefer ranking.

## How
- Pin the pdbmux Deployment image to `v0.2.0`.
- Leave `pdbmux-env` unchanged: `PDBMUX_LISTEN`, `PDBMUX_BACKENDS`, `PDBMUX_MERGE` are the only keys v0.2.0 reads from this ConfigMap, and `PDBMUX_PRIMARY`/`PDBMUX_PREFER` are already gone (#452).

Reviewed-on: #455
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 17:05:45 +10:00
unkin-agent 2fcb3d70d5 Bump artifactapi to v3.11.2 (#453)
Why: pick up v3.11.2, which moves DB migrations onto golib/pg with no behavior change.

- Bump the artifactapi API and UI image pins from v3.11.1 to v3.11.2.

Reviewed-on: #453
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 16:04:58 +10:00
unkin-agent 1172aa3e96 pdbmux: order backends new-first and drop primary/prefer (#452)
## Why
pdbmux#9 drops primary/prefer and makes configured backend order the only tie-break, so the current `old`-first list would silently reverse which PuppetDB wins.

## How
- Order `PDBMUX_BACKENDS` with `new=http://puppetdb.puppet.svc.cluster.local:8080` first and `old=http://puppetdbapi.service.consul:8080` second, URLs unchanged.
- Drop `PDBMUX_PRIMARY` and `PDBMUX_PREFER`; both already resolve to `new`, on the deployed v0.1.0 image (estate defaults) and on pdbmux main (first-backend fallback), so the rendered behaviour is unchanged today.
- Refresh the configmap and deployment comments to describe order-based precedence.

Merge this before pdbmux#9.

Reviewed-on: #452
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 16:04:40 +10:00
unkin-agent e24b19412a woodpecker: add vimpack-ci ServiceAccount (#454)
Why: the new vimpack repo's woodpecker pipelines set `serviceAccountName: vimpack-ci`, which does not exist in the woodpecker namespace yet.

- Add bare `vimpack-ci` ServiceAccount in `apps/base/woodpecker/` and register it in the kustomization.

Reviewed-on: #454
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 13:42:33 +10:00
unkin-agent b2b82e6a6c puppet: use OPENVOXSERVER_JAVA_ARGS for puppetserver JVM args (#451)
## Why

The `ghcr.io/openvoxproject/openvoxserver` image reads `OPENVOXSERVER_JAVA_ARGS` (`/etc/default/puppetserver`: `JAVA_ARGS=$OPENVOXSERVER_JAVA_ARGS`) and ships no `PUPPETSERVER_*` alias, so our heap/JMX flags have been inert since the fork switch — live masters and compilers run the image default `-Xms1024m -Xmx1024m` with no JMX.

## How

- Rename `PUPPETSERVER_JAVA_ARGS` to `OPENVOXSERVER_JAVA_ARGS` in `puppetserver-master-config`, `puppetserver-compiler-config` and `puppetserver-init-config`.
- Rename the same var on the `perms-and-dirs` init container in `deployment_puppetserver-compiler.yaml`.
- Flag values are unchanged (`-Xms1024m -Xmx3072m` plus the JMX flags); heap tuning is a separate call.

Reviewed-on: #451
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:30:22 +10:00
unkin-agent 7f15488332 Bump encapi image to v0.1.2 (#450)
Why: pick up encapi v0.1.2, which moves DB migrations onto golib/pg with no behavior change (release pipeline green).

- Bump the encapi deployment container image from v0.1.1 to v0.1.2

Reviewed-on: #450
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:23:45 +10:00
unkin-agent 8ea2925561 Set puppetserver environment timeout to 0 (#449)
## Why

An unlimited environment timeout serves stale catalogs after code changes.

## How

- Set `OPENVOXSERVER_ENVIRONMENT_TIMEOUT: "0"` in the puppetserver master and compiler ConfigMaps.

Reviewed-on: #449
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:23:04 +10:00
unkin-agent 236563a37f Bump arrproxy images to v0.6.1 (#448)
Pick up the golib/pg migration runner refactor released in arrproxy v0.6.1; no behavior change.

- Bump arrproxy-api and arrproxy-ui image tags v0.6.0 -> v0.6.1

Reviewed-on: #448
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:15:47 +10:00
unkin-agent df38b6b75f Add golib-ci ServiceAccount for woodpecker (#447)
## Why

The new `golib` repo's Woodpecker pipeline needs a dedicated ServiceAccount to run its CI steps under.

## How

- Add bare ServiceAccount `golib-ci` in the `woodpecker` namespace and wire it into the base kustomization.

Reviewed-on: #447
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-31 06:58:45 +10:00
unkin-agent c5fff07643 Bump repospawner to v0.1.1 (#446)
Why: repospawner v0.1.1 stops spawned job pods from automounting the API token.

- Bump the deployment image and the matching REPOSPAWNER_IMAGE env value to v0.1.1

Reviewed-on: #446
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-30 17:07:58 +10:00
unkin-agent a9a66a07b1 Deploy repospawner v0.1.0 (#445)
## Why

repospawner v0.1.0 is built and its Vault kubernetes auth role is applied, but nothing deploys it. It turns a "I want a new repository" request into a terraform-git pull request, follows that PR to merge, and optionally activates the repo in Woodpecker, so the review gate stays where it is instead of moving into an agent's hands.

## How

- Add `apps/base/repospawner/`: namespace, ServiceAccount `repospawner`, `default` VaultAuth for VSO, and a namespaced Role/RoleBinding granting jobs create/get/list/watch/delete plus pods and pods/log reads (mirrors mediamover).
- Deployment pinned to `artifactapi.k8s.syd1.au.unkin.net/docker-internal/repospawner:v0.1.0`, one replica with the `Recreate` strategy because request state is in memory and rebuilt from Job labels; the same image reference is passed down as `REPOSPAWNER_IMAGE` so the spawned Jobs stay in step.
- Mount a projected `audience: vault` service account token at `/var/run/secrets/vault` — the app logs into Vault natively rather than through VSO — and the `repospawner-woodpecker` Secret at `/etc/repospawner/woodpecker`, optional so the server still starts and refuses `woodpecker: true` with 503 when it is absent.
- Two VaultStaticSecrets: `oauth-credentials` from `kv/kubernetes/namespace/repospawner/default/oauth-credentials` and `repospawner-woodpecker` (key `token`) from `.../default/woodpecker`, with reloader annotations on both consumers.
- oauth2-proxy front door on the watchstate/mediamark pattern, gated on `akP-repospawner-admin` via the `ak_groups` claim and re-checked by the app from `X-Forwarded-Groups`; public `repospawner.unkin.net` on the reflected wildcard and internal `repospawner.k8s.syd1.au.unkin.net` on `vault-issuer`, both routed to the oauth2 Service.
- Register the overlay in the platform ApplicationSet and AppProject, and append `repospawner` to the wildcard Certificate's two reflector namespace lists.

Depends on the terraform-authentik `repospawner` client being applied and `kv/kubernetes/namespace/repospawner/default/oauth-credentials` + `.../woodpecker` being seeded.

Reviewed-on: #445
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-30 15:40:06 +10:00
unkin-agent abf73bb5bf arrproxy: v0.6.0 self-migrating, drop external migrate Job (#444)
## Why

arrproxy v0.6.0 applies its own schema at startup under a Postgres advisory lock and holds `/readyz` until the schema is current, so every replica is safe to roll without an external gate. The wave-1 psql `arrproxy-migrate` Job and its SQL ConfigMap now only re-run idempotent statements the app already owns — dead weight, a second source of truth for the schema, and a standing drift trap whenever the app's embedded migrations move ahead of the manifests.

## How

- Bump `arrproxy-api` and `arrproxy-ui` to `v0.6.0`.
- Delete `migrate-job.yaml` and `migrations-configmap.yaml` and drop both from the arrproxy kustomization.
- Keep the wave-0/wave-2 split: wave 2 still orders the api behind the wave-0 CNPG Cluster and VSO-synced Secrets, which is independent of the migrate Job; the stale "serve only after the wave-1 migrate Job" comment is corrected.
- Rendered diff vs `main` is exactly the two image bumps plus the `arrproxy-migrate` Job and `arrproxy-migrations` ConfigMap disappearing; `kustomize build --enable-helm apps/overlays/au-syd1/arrstack` and pre-commit both clean.

Reviewed-on: #444
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-30 15:30:43 +10:00
unkin-agent b355d6aafb mediamark: deploy the media marking app (#441)
## Why

mediamark needs a home in the cluster: it marks/organises media on the shared mediastore tree and drives the adult-tier sonarr/radarr through arrproxy's hash routes. This adds the full app to the `media` project, mirroring the watchstate two-host oauth2-proxy pattern and the arrstack static-CephFS + projected-API-keys patterns.

## How

- Adds `apps/base/mediamark/`: namespace, VaultAuth (`k8s/au/syd1`, role `default`), three VaultStaticSecrets, static mediastore PV/PVC, the app Deployment, oauth2-proxy ConfigMap/Deployment, two Services, and internal + external Gateway/HTTPRoute pairs.
- Binds a dedicated static PV (`mediamark-mediastore`, own `volumeHandle`, `claimRef`-pinned) to the same CephFS mediastore subvolume arrstack/fafflix/cheeztv use, RWX 10Ti Retain, mounted at `/media`.
- Runs the app as 1000:1000 (deliberately not 65532) so it owns files on the shared media tree and hardlink/rename moves stay valid; read-only root filesystem, all caps dropped, no service-account token, `/livez` + `/readyz` probes.
- Projects the sonarr/radarr API keys as one file per app under `/etc/mediamark/keys`, mirroring arrproxy's keys projection, with reloader annotations on both secrets.
- Fronts both `mediamark.unkin.net` (traefik-external, reflected Let's Encrypt wildcard, no cert-manager annotations) and `mediamark.k8s.syd1.au.unkin.net` (traefik-internal, vault-issuer) with a single oauth2-proxy using a relative `/oauth2/callback` redirect; gated on `akP-mediamark-user` and passing identity to the app as `X-Forwarded-Groups` via `PASS_USER_HEADERS`.
- Appends `mediamark` to the `wildcard-unkin-net` Certificate's two reflector namespace lists, and registers the app in `argocd/applicationsets/media.yaml` + `argocd/projects/media.yaml` with a passthrough `apps/overlays/au-syd1/mediamark` overlay.

## Prerequisite seeds (Ben, before pods go Ready)

These KV paths must exist under `kv/kubernetes/namespace/mediamark/default/` — the `mediamark/default` templated policy already grants read, so no terraform-vault change is needed:

- `oauth-credentials` — needs `client_id` and `cookie_secret` added alongside the existing `client_secret` (Authentik mediamark provider; both absolute callback URIs registered there).
- `sonarr` — key `apitoken`.
- `radarr` — key `apitoken`.

## Validation

- `kustomize build --enable-helm apps/overlays/au-syd1/mediamark` (18 resources) and `.../cert-manager` both build.
- kubeconform clean on both touched overlays.
- `pre-commit run --all-files` passes.

Reviewed-on: #441
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-30 15:07:41 +10:00
unkin-agent 221c575a44 arrproxy: bump images to v0.5.0 (per-token method scoping) (#443)
## Why

arrproxy v0.5.0 ships per-token HTTP method scoping for machine tokens, so a minted token can be limited to e.g. `GET` only. Zero-downtime: the mint-API field is additive and existing tokens get an empty methods list, which means unrestricted — they behave exactly as before.

## How

- Bump `arrproxy-api` and `arrproxy-ui` pins from v0.4.0 to v0.5.0.
- Mirror repo migrations `0002_tier_tokens.sql` and `0003_token_methods.sql` into the migrations ConfigMap. It had drifted at 0001 while v0.4.0 already queried `tier`/`read_only`, and every v0.5.0 token query selects `methods` — without this the new API errors on every token read.
- Have the wave-1 migrate Job apply all three files in order. Every statement is `IF NOT EXISTS`, so a resync over an already-migrated database is a no-op.

Rendered `kustomize build --enable-helm apps/overlays/au-syd1/arrstack` diff vs main is exactly the two image tags, the two added ConfigMap keys, and the two added `-f` args.

Reviewed-on: #443
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-30 14:23:44 +10:00
unkin-agent 1ba6180e96 woodpecker: add repospawner-ci ServiceAccount (#442)
The new repospawner repo needs a Woodpecker CI pipeline, and every pipeline step must run under its own ServiceAccount in the woodpecker namespace.

- Add `apps/base/woodpecker/serviceaccount_repospawner_ci.yaml` (bare SA `repospawner-ci`, ns woodpecker), mirroring the existing mediamark-ci SA.
- Register it in the woodpecker kustomization resources list.

Reviewed-on: #442
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-30 14:21:42 +10:00
unkin-agent d1085f0ae2 logging: use canonical upstream image names (#433)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit (logarchiver/logviewer are untouched).

Changes:
- Point the ClickHouseInstallation and the clickhouse-schema job at `docker.io/clickhouse/clickhouse-server:24.8`.
- Point the logviewer oauth2-proxy cert-combine init container at `docker.io/library/alpine:3`.
- Point the NATS bootstrap job at `docker.io/natsio/nats-box:0.18.0`.
- Point the NATS chart values at `docker.io/library/nats` and `docker.io/natsio/nats-server-config-reloader`.
- Point all three Vector values files (agent, aggregator, vm-ingest) at `docker.io/timberio/vector`.
- Drop the now-wrong "pulled through the artifactapi dockerhub remote" comments in the NATS and vector-agent values.

Tags/digests unchanged and the `repository`/`tag` split is preserved. `kustomize build --enable-helm apps/overlays/au-syd1/logging` differs from main only in those nine image strings.

Extra found, not changed here: `.woodpecker/vector-test.yaml` still pins its CI step image to `artifactapi.k8s.syd1.au.unkin.net/dockerhub/timberio/vector:0.57.0-debian`. That is a Woodpecker step image rather than a namespace manifest, so it is left out to keep this PR to the logging namespace — say the word and I will fix it separately.

Reviewed-on: #433
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:55:47 +10:00
unkin-agent e4d93ef4fe valkey-operator-system: use canonical ghcr.io registry (#437)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Set the valkey-operator chart `image.registry` to `ghcr.io`.

The `registry`/`repository`/`tag` split is untouched otherwise, so the rendered image is `ghcr.io/valkey-io/valkey-operator:v0.5.0`. `kustomize build --enable-helm apps/overlays/au-syd1/valkey-operator-system` differs from main only in that image string. No other proxied image refs in the file (the `helmCharts[].repo` entry in kustomization.yaml is a Helm chart repo, not a container registry, so it stays on artifactapi).

Reviewed-on: #437
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:52:52 +10:00
unkin-agent 1169d796e7 grafana: stop pinning the internal CA for Authentik OAuth (#440)
## Why

`identity.unkin.net` moved from an internal `unkin.net` CA-issued cert to the LetsEncrypt `*.unkin.net` wildcard. `auth.generic_oauth`'s `tls_client_ca` pointed Grafana at the internal root only, so the OAuth handshake to the LE-issued cert now fails. Grafana's image trust store already contains the public roots.

## Changes

- Remove `tls_client_ca: /etc/grafana/vault-ca/ca.crt` (and its stale comment) from `auth.generic_oauth`.
- Remove the now-unused `vault-ca-cert` volume and volumeMount from the Grafana pod spec — nothing else in the pod referenced it (the CNPG `endpointCA` reference to `vault-ca-cert` for `s3.ceph.unkin.net` is a separate resource and stays).
- Leave the auth/token/api URLs, scopes and `role_attribute_path` untouched.

Reviewed-on: #440
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:52:30 +10:00
unkin-agent aedb721b3e argocd: drop internal-CA rootCA pin from Authentik oidc.config (#439)
## Why

`identity.unkin.net` moved from an internal `unkin.net` CA-issued cert to the LetsEncrypt `*.unkin.net` wildcard. The `oidc.config` block pinned the internal root as the *only* trust anchor, so argocd-server now rejects OIDC discovery with `x509: certificate signed by unknown authority` and SSO login is broken. The stock image trust store already carries the public roots.

## Changes

- Remove the `rootCA:` block from `argocd-cm`'s `oidc.config` for the `https://identity.unkin.net/application/o/argocd/` issuer.
- Replace the now-false internal-CA rationale comment with a one-liner noting the LE-issued cert needs no pin.
- Leave issuer, clientID, clientSecret ref, `requestedScopes` (incl. `ak_groups`) and `requestedIDTokenClaims` untouched.

Reviewed-on: #439
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:52:12 +10:00
unkin-agent 74ad2c8773 woodpecker: add mediamark-ci service account (#438)
The mediamark Woodpecker docker step needs a dedicated ServiceAccount so it can push to the trusted in-cluster registry, mirroring the existing arrproxy-ci setup.

- Add bare ServiceAccount `mediamark-ci` in namespace `woodpecker` and register it in the woodpecker base kustomization.

Reviewed-on: #438
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:21:21 +10:00
unkin-agent 6b5b129ad6 clickhouse-system: use canonical upstream image names (#436)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Point the Altinity chart crdHook at `docker.io/bitnami/kubectl`.
- Point the operator at `docker.io/altinity/clickhouse-operator`.
- Point the metrics exporter at `docker.io/altinity/metrics-exporter`.
- Reword the header comment that claimed all images are pulled through the artifactapi dockerhub remote.

Only the `repository` keys change; the chart still supplies the tags (0.27.2 / latest), so rendered tags are identical. `kustomize build --enable-helm apps/overlays/au-syd1/clickhouse-system` differs from main only in those three image strings. No other proxied refs in the file.

Reviewed-on: #436
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:20:10 +10:00
unkin-agent 38a41bd44b watchstate: use canonical upstream image name for alpine (#435)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Point the watchstate oauth2-proxy cert-combine init container at `docker.io/library/alpine:3`.

Tag unchanged. `kustomize build --enable-helm apps/overlays/au-syd1/watchstate` differs from main only in that image string. No extra proxied refs in the file (the oauth2-proxy image itself is already canonical `quay.io/...`).

Reviewed-on: #435
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:19:36 +10:00
unkin-agent da1d812eec netbox: use canonical upstream image name for redis_exporter (#434)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Point the netbox valkey metrics sidecar at `docker.io/oliver006/redis_exporter:v1.89.0`.

Tag unchanged. `kustomize build --enable-helm apps/overlays/au-syd1/netbox` differs from main only in that image string. No extra proxied refs in the file (the `valkey/valkey:8-alpine` container is already a bare upstream name).

Reviewed-on: #434
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:19:27 +10:00
unkin-agent df89947f47 litellm: use canonical upstream image name for redis_exporter (#432)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Point the litellm redis metrics sidecar at `docker.io/oliver006/redis_exporter:v1.89.0`.

Tag unchanged. `kustomize build --enable-helm apps/overlays/au-syd1/litellm` differs from main only in that image string. No extra proxied refs in the file (the `redis:7-alpine` container is already a bare upstream name).

Reviewed-on: #432
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:19:01 +10:00
unkin-agent 7f928dddfc gitea: use canonical upstream image name for redis_exporter (#431)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Point the gitea valkey metrics sidecar at `docker.io/oliver006/redis_exporter:v1.89.0`.

Tag unchanged. `kustomize build --enable-helm apps/overlays/au-syd1/gitea` differs from main only in that image string. No extra proxied refs in the file (the `valkey/valkey:8-alpine` container is already a bare upstream name).

Reviewed-on: #431
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:18:44 +10:00
unkin-agent b459e9a90a authentik: use canonical upstream image name for redis_exporter (#430)
rke2's `registries.yaml` already rewrites upstream image names to the artifactapi mirror, so manifests must carry canonical upstream names. Only in-house `artifactapi.k8s.syd1.au.unkin.net/docker-internal/...` images stay explicit.

Changes:
- Point the authentik redis metrics sidecar at `docker.io/oliver006/redis_exporter:v1.89.0`.

Tag unchanged. `kustomize build --enable-helm apps/overlays/au-syd1/authentik` differs from main only in that image string. No extra proxied refs in the file (the `redis:7-alpine` container is already a bare upstream name).

Reviewed-on: #430
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-08-29 21:18:33 +10:00