Commit Graph

482 Commits

Author SHA1 Message Date
unkin-agent f42ca99a8c Redirect vlogs root to the VictoriaLogs UI
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
Add an Exact / rule on the vlogs HTTPRoute that 302s to /select/vmui/,
so the bare host lands on the UI instead of the index page.
2026-09-27 17:45:54 +10:00
unkin-agent 216b1d72ac Split bind-internal DNSRecords into one file per record (#501)
records.yaml had grown to ten DNSRecord documents across two zones, so finding or reviewing a single record meant scanning the whole file and every change touched it.

- move each record to authoritative/<zone>/<type>/<record>.yaml
- add a kustomization.yaml per zone and per type directory
- keep the git.unkin.net cutover record commented out, alongside a commented kustomization entry
- move the DNSRecord-namespace rationale onto the authoritative kustomization
- delete records.yaml

Rendered output is unchanged.

Reviewed-on: #501
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 17:43:16 +10:00
unkin-agent 7581578df1 Split bind-external DNSRecords into one file per record (#500)
A single records.yaml holding every DNSRecord means any record change touches a shared file, and deleting one record is a hunk edit rather than a file removal. One file per record under <zone>/<type>/<record>.yaml makes each record independently editable and removable.

- move acme-apex-ns to acme-unkin-net/ns/apex.yaml and acme-ns1-a to acme-unkin-net/a/ns1.yaml
- add kustomization.yaml per zone and type directory, referencing directories from the parent
- carry the delegation rationale onto the zone kustomization
- drop records.yaml and reference the zone directory from the app kustomization

Rendered output is unchanged.

Reviewed-on: #500
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 17:28:37 +10:00
unkin-agent 6515e7f637 Expose VictoriaLogs UI at vlogs.unkin.net behind oauth2-proxy (#498)
The vlselect query UI is only reachable in-cluster, so every log search needs a port-forward. Publish it at vlogs.unkin.net behind oauth2-proxy on the external (DMZ) Traefik.

- Add `apps/base/logging/vlogs`: external Gateway, http->https redirect and main HTTPRoute
- Terminate TLS with the reflected Let's Encrypt `*.unkin.net` wildcard; reflect it into `logging`
- Publish the `vlogs` A record at the DMZ gateway VIP from the bind-operator `unkin.net` zone
- Route all traffic through the `vlogs-oauth2` Service, upstreaming to `vlselect-logs:9471`
- Gate on the Authentik `vlogs` application, group `akP-vlogs-admin`
- Read OIDC credentials from `kv/kubernetes/namespace/logging/default/vlogs-oauth-credentials`

Depends on unkin/terraform-authentik#40.

Reviewed-on: #498
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 17:21:40 +10:00
unkin-agent 90d512fadf Drop the gocache serve Deployment and nginx sidecar (#499)
`connect` mode gives the client no local cache, so every cache operation is a network round trip -- about 2.7x slower than pointing the plugin straight at S3. Developers and CI run the plugin in direct mode instead.

- Remove the `go-cache-plugin serve` Deployment and its nginx stream-proxy ConfigMap
- Remove the PureLB Service, freeing `198.18.200.11`
- Keep the `gocache` bucket and `gocache-s3` Secret for direct-mode clients

Reviewed-on: #499
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 11:20:59 +10:00
unkin-agent 757ae5b240 Add a VictoriaLogs cluster and point logs-ingest at it (#488)
The k8s log pipeline stores to ClickHouse via NATS+vector, while the VM estate ships journald to a separate puppet-managed VictoriaLogs cluster. Consolidating on VictoriaLogs in-cluster collapses the two paths, and the logs-ingest gateway has no clients yet so it can be repointed now, ahead of the puppet change.

- add VLCluster `logs` at v1.52.0 (2 vlinsert, 2 vlselect, 3 vlstorage, 180d retention, 250Gi each on cephrbd-fast-delete)
- cap vlstorage disk use at 220GiB per node so 180d stays time-based rather than disk-bound
- repoint the logs-ingest HTTPRoute at `vlinsert-logs:9481`
- add a VictoriaLogs Grafana datasource and install its plugin

Nothing is removed here; NATS, ClickHouse, vector, logarchiver and logviewer keep running until a follow-up drops them.

Reviewed-on: #488
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:33:58 +10:00
unkin-agent 02f877540c Add gocache serve Deployment with nginx stream sidecar (#496)
`go-cache-plugin serve` binds `127.0.0.1` only, so nothing outside the pod can reach it and laptops have no way to use the S3-backed Go cache without holding RGW credentials.

- Run `go-cache-plugin serve` against the `gocache` bucket, path-style, explicit region to skip the GetBucketLocation probe
- Add an nginx sidecar stream-proxying `:9090` to the loopback plugin port, `proxy_timeout 2h`
- Publish it on PureLB `198.18.200.11`, `externalTrafficPolicy: Local` so the client IP reaches the allow rules
- Restrict to workstation + pod CIDRs: GOCACHEPROG is unauthenticated and a poisoned entry runs in every consuming build

Merge only after `docker-internal/go-cache-plugin:v0.1.0` is published.

Reviewed-on: #496
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:32:32 +10:00
unkin-agent 148dac8ca2 Declare the acme.unkin.net nameservers (#495)
The zone was seeded with an apex `NS ns1.acme.unkin.net` glued to the primary pod IP. Both were later corrected by hand, so the live RRset and the ns1 address exist only in the zone journal -- a reseed republishes the pod IP and breaks DNS-01 for every `*.unkin.net` cert. Declaring them makes git the source of truth.

- declare the two published apex NS names
- declare the in-zone ns1 address, which a seed would otherwise glue to the pod IP

Matches what the zone serves today, so applying it changes no records. Requires bind-operator v0.3.0.

Reviewed-on: #495
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:10:13 +10:00
unkin-agent 53e5846c18 Roll bind-operator to v0.3.0 (#494)
v0.3.0 converges a zone apex NS onto its declared nameservers instead of leaving the seed placeholder, which publishes a primary pod IP. The CRD moves with the image because the operator reads the new `spec.nameservers` field.

- pin the bind-operator image to v0.3.0
- pull the CRDs from the v0.3.0 tag

Reviewed-on: #494
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:06:21 +10:00
unkin-agent 9fed5decc8 bump victoria-metrics-operator chart to 0.67.3 (#493)
The vm-system overlay pins victoria-metrics-operator chart 0.57.1 (operator v0.66.1), eight operator minors behind upstream, so the cluster runs without newer CRD fields and reconciler fixes.

- bump the victoria-metrics-operator helmChart version to 0.67.3 (operator v0.74.1)

Reviewed-on: #493
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:05:47 +10:00
unkin-agent 20077f1029 Move the haproxy edge behind the external Traefik (#492)
The haproxy edge holds its own DMZ VIP, a second public entry point alongside
traefik-external that must be firewalled and DNS'd separately. Traefik can
front it with TLS passthrough, leaving haproxy's certs and backends untouched.

- Add a `traefik-external` Gateway: HTTP :80 plus Passthrough TLS :443.
- TLSRoute the 12 `fe_https.map` hostnames to haproxy:443; HTTPRoute 301s :80.
- Make the Service ClusterIP on 443 only, releasing 198.18.199.1.
- Drop `fe_http`, `be_letsencrypt` and `fe_http.map`; certs are DNS-01 only.

Client IP now reads as a Traefik pod — the Gateway provider cannot emit PROXY protocol to a TLSRoute backend. `sessionAffinity` goes too (it would pin Traefik pods, not clients); SRVNAME cookies keep persistence.

Reviewed-on: #492
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 23:38:07 +10:00
unkin-agent 426a399f31 Drop stalwart mail proxying from the haproxy edge (#491)
Stalwart was only ever a test deployment. The daemon is dead on all three
backend VMs and nothing public depends on it — `unkin.net` MX points at Google —
so the edge is proxying mail to nowhere and the tcp frontends make `defaults`
emit 20 spurious HTTP-mode warnings.

- Drop the `fe_smtp`, `fe_submission`, `fe_imap` and `fe_imaps` frontends.
- Drop the five `be_stalwart_*` backends and their map entries in `fe_http.map`/`fe_https.map`.
- Drop the now-unused 25/143/587/993 Service and container ports.

`haproxy -c` on the rendered config: exit 0, 0 warnings (was 20), 0 alerts.

Reviewed-on: #491
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 21:25:49 +10:00
unkin-agent 5341253573 Add Ceph RGW bucket for the shared Go build cache (#489)
Go builds on CI and laptops each rebuild the same packages from scratch. A
GOCACHEPROG backend needs an S3 bucket plus credentials before anything can
point at it, so provision those first. The bucket lives in the woodpecker
namespace because CI is the primary consumer and reads the Secret there.

- add Bucket and ObjectStoreUser for the shared Go build cache
- use default (replicated) placement rather than the ec target, since a build
  cache is millions of small objects
- purge and drop the bucket and user on delete; the cache is disposable

Nothing consumes the bucket yet.

Reviewed-on: #489
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:56:51 +10:00
unkin-agent d48125d699 Add job and start deadlines to the g10k-code CronJob (#490)
A g10k-code job wedged in ContainerCreating on a failed CephFS mount and never reached a terminal condition, so it stayed in the CronJob active list and `concurrencyPolicy: Forbid` skipped every following minute. No Puppet code reached the estate for 6 days, and the piled-up missed slots crossed the controller 100-slot cap into `TooManyMissedTimes`. The CronJob carried no deadlines at all.

- Cap a job at `activeDeadlineSeconds: 300` on the Job spec, so a hang is failed as `DeadlineExceeded` and drops out of the active list (healthy runs take 16-18s).
- Set `startingDeadlineSeconds: 200`, bounding missed-schedule look-back to ~3 slots so the count cannot reach 100.

Reviewed-on: #490
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:55:32 +10:00
unkin-agent abf6bfae88 Drop dead X-Frame-Options rules from the haproxy edge (#487)
The 13 `X-Frame-Options DENY if acl_<host>` rules in `fe_https` have never fired:
their ACLs use `req.hdr(host)`, a request-direction fetch that is invalid in a
response ruleset, so HAProxy rejects them at config-check time. Carried over
verbatim from the Puppet LXD config during the k8s move.

- Remove the 13 dead `http-response set-header X-Frame-Options` rules.
- Remove the 15 now-orphaned `acl acl_*` definition lines.

Not switching the header on: it has never been live, and Grafana/Gitea send their
own. `haproxy -c` warnings drop 33 -> 20; the two working response headers stay.

Reviewed-on: #487
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:36:26 +10:00
unkin-agent 4190785389 Move the au-syd1 haproxy edge into Kubernetes (#485)
The au-syd1 edge proxy runs on a hand-managed LXD container outside the cluster, with no HA and no shared config source.

- Add `apps/base/haproxy/`: 3 replicas behind the DMZ LoadBalancer 198.18.199.1, config from a ConfigMap, wildcard certs from reflected secrets.
- Keep source IPs via `externalTrafficPolicy: Local`; `sessionAffinity: ClientIP` stands in for the stick-table peers a Deployment cannot name.
- Drain on shutdown: `hard-stop-after 2m`, a preStop SIGUSR1 soft-stop, 150s grace.
- Bind the stats listener to 127.0.0.1 so it is port-forward only.
- Register the app in the platform project and ApplicationSet.

Reviewed-on: #485
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 18:22:57 +10:00
unkin-agent 812a9a2f2b Publish real delegation records for acme.unkin.net (#486)
The acme.unkin.net zone still serves only the bind-operator seed apex: NS ns1.acme.unkin.net glued to A 10.42.6.38, a pod IP no pod holds. The parent delegates to acme-ns1.unkin.net, but public resolvers have already promoted the child NS RRset, so when the cached address expires DNS-01 fails for every unkin.net wildcard at once.

- Add apex NS acme-ns1.unkin.net., matching the parent delegation (out of zone, no glue needed).
- Point ns1.acme.unkin.net at 103.216.191.185 so resolvers holding the seeded NS name still reach the zone.
- The operator seed placeholder itself is tracked separately in bind-operator.

---------

Co-authored-by: unkin-agent <agent@unkin.net>
Reviewed-on: #486
Co-authored-by: Unkin Agent <unkin-agent@unkin.net>
Co-committed-by: Unkin Agent <unkin-agent@unkin.net>
2026-09-26 16:42:02 +10:00
unkin-agent f37749523d Add *.main and *.ceph wildcard certificates for haproxy (#484)
The haproxy edge terminates TLS for hosts under `main.unkin.net` and `ceph.unkin.net`, which the single `*.unkin.net` wildcard does not cover.

- Add cert-manager Certificates for both wildcards from the `letsencrypt` ClusterIssuer.
- Reflect the minted secrets into the `haproxy` namespace.

Needs these records in the public unkin.net zone first:
`_acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.`
`_acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.`

Reviewed-on: #484
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 16:38:31 +10:00
unkin-agent fe51aa07be Give the puppetserver compilers the Vault cert helpers (#482)
profiles::pki::vault and profiles::ssh::sign shell out to
/usr/local/bin/certmanager and /usr/local/bin/sshsignhost from generate()
during catalog compilation. Neither binary exists in the compiler image, so
every node using them fails to compile.

- install certmanager v0.2.0 and sshsignhost v0.1.0 onto the shared bin volume with sha256 verification
- wrap both at /usr/local/bin from a pre-default entrypoint hook, failing startup loudly if either is missing
- mount read-only Vault configs for both: kubernetes auth on k8s/au/syd1, internal CA verified rather than skipped

Reviewed-on: #482
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-24 21:15:09 +10:00
unkin-agent cdaab736b5 Bump jellyfin-ha to v0.4.0 (#483)
The deployed v0.3.3 build returns 500 from /Shows/NextUp on PostgreSQL, breaking the home screen, and lets replicas diverge: library-visibility and shared-config changes never propagate, user data (resume, played state, favourites, ratings) is overwritten between pods, and eight scheduled tasks run on every replica instead of only the scan leader. v0.4.0 carries the fixes.

- Pin cheeztv and fafflix to jellyfin-ha:v0.4.0.

No config change needed: cross-pod invalidation reuses the transcode-store Redis connection string both apps already set.

Reviewed-on: #483
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-22 22:57:06 +10:00
unkin-agent b31517e6d9 Merge pull request #481 from benvin/jellyfin-sso-valkey-state
Roll jellyfin-ha to v0.3.3 and drop Service session affinity

Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 14:02:48 +10:00
unkin-agent 5a74b2cec6 Roll bind-operator to v0.2.7 (journal-aware zone seeding) (#480)
Deploy bind-operator v0.2.7. The operator seeded a fresh skeleton zone file at serial 1 over zones whose BIND journal was still on disk at a higher serial; BIND rejected the inconsistent pair (`addzone failed: out of range`) and, with a PVC per replica, the stale journal outlived restarts while every reconcile rewrote the skeleton, so it never converged. That SERVFAILed roughly 1 in 3 authoritative answers for k8s.syd1.au.unkin.net and resolvers cached the failures.

- Bumps the operator image to v0.2.7
- Bumps the CRD install pin to the v0.2.7 tag, which changes the CRDs

Expect one rolling restart of the operator Deployment as the new image lands.

Reviewed-on: #480
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 00:53:47 +10:00
unkin-agent f14bcc4d2a Add woodpecker ServiceAccount for jellyfin-plugin-sso CI (#479)
The new `unkin/jellyfin-plugin-sso` fork is getting a Woodpecker pipeline, and its build step will set `serviceAccountName: jellyfin-plugin-sso`. Without the SA declared here the pipeline pods fail to schedule.

- add a bare ServiceAccount `jellyfin-plugin-sso` in the `woodpecker` namespace
- register it in the woodpecker base kustomization

The step only builds .NET code, so no Vault kube-auth role or RBAC is needed.

Reviewed-on: #479
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 17:32:31 +10:00
unkin-agent b01e4c3241 Rewrite slash-less Authentik token endpoint to the canonical path (#478)
Authentik advertises the token endpoint with a trailing slash, but some OIDC clients (the ArgoCD iOS app) POST to /application/o/token without one; Django's APPEND_SLASH will not redirect a POST, so the token exchange gets 405 and login fails.

- Add an exact-match rule on /application/o/token to the authentik and authentik-internal HTTPRoutes.
- Rewrite it to /application/o/token/ with a URLRewrite ReplaceFullPath filter, preserving the method and the authentik-server backend.
- Leave the catch-all PathPrefix rule untouched; exact matches outrank it in Gateway API precedence.

Reviewed-on: #478
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 16:10:30 +10:00
unkin-agent bbd5bdaa95 Enable PKCE for ArgoCD OIDC login (#477)
The Authentik client for ArgoCD is now public (the iOS app can't hold
a secret), so Authentik no longer enforces client_secret on token
exchange. PKCE replaces that as the protection against
authorization-code interception.

- Add `enablePKCEAuthentication: true` to the `oidc.config` block in
  `argocd-cm-patch.yaml`
- Note why PKCE is needed now that the client is public

Reviewed-on: #477
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 16:10:05 +10:00
unkin-agent 4762cf9e03 Add VMPodScrape for authentik-server metrics (#476)
Authentik server pods expose django_prometheus metrics on :9300, but only ldap-outpost and redis-exporter are scraped in this namespace. Add the missing per-app scrape.

- add apps/base/authentik/server-vmpodscrape.yaml selecting app.kubernetes.io/name=authentik, component=server on the metrics port
- wire it into apps/base/authentik/kustomization.yaml

Reviewed-on: #476
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 15:34:26 +10:00
unkin-agent c83a886e74 Enable pg_stat_statements on the authentik postgres cluster (#475)
The cluster preloads no statement-statistics library, so there is no per-query cost attribution in postgres and slow query paths have to be inferred from application-side metrics instead of read straight out of the database.

- preload `pg_stat_statements`
- set `pg_stat_statements.max` and `.track`, which is what makes CNPG manage the extension and create it in every database

Requires a postgres restart. Stacked on `benvin/authentik-cnpg-resources`.

Reviewed-on: #475
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:48:45 +10:00
unkin-agent 9a7200636c Raise authentik postgres CPU, memory and buffer sizing (#474)
The 500m CPU limit is a 50ms CFS quota per 100ms period, so the postgres pods are throttled on every burst even at ~0.01 cores average and each query pays that latency. 128MB of `shared_buffers` and a 256MB cache estimate also under-serve the planner on the joins authentik issues on its hot read paths.

- raise resources to requests `500m`/`1Gi`, limits `2`/`2Gi`
- raise `shared_buffers` to 512MB and `effective_cache_size` to 1536MB
- hold the post-incident memory headroom multiple over `shared_buffers`

Rolling restart with switchover. Stacked on `benvin/authentik-hot-standby-feedback`.

Reviewed-on: #474
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:35:16 +10:00
unkin-agent 9535bad9bc Enable hot_standby_feedback on the authentik postgres cluster (#473)
Authentik serves multi-second API reads from the CNPG hot standbys. Those reads outlive `max_standby_streaming_delay`, so recovery cancels them with `canceling statement due to conflict with recovery`, which authentik surfaces as HTTP 500 — enough to break a terraform apply mid-run.

- set `hot_standby_feedback` on so replicas report their oldest xmin to the primary and long reads stop being cancelled
- SIGHUP reload only, no restart or switchover
- retained-dead-tuple cost is negligible on a ~155MB database

Reviewed-on: #473
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:34:33 +10:00
unkin-agent 34dd70435e Auto-reload cheeztv and fafflix on plugin ConfigMap change (#471)
Edits to the cheeztv/fafflix plugin ConfigMaps only reach the pods via the inject-plugin-config initContainer, so a config change sat inert until someone manually rolled the StatefulSet. Reloader is deployed cluster-wide with autoReloadAll disabled, so each workload has to opt in.

- annotate both StatefulSets with configmap.reloader.stakater.com/auto: "true"

Reviewed-on: #471
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 12:25:08 +10:00
unkin-agent 6cc752336e Point jellyfin SSO at public Authentik hostname (#470)
The internal-CA identity.k8s.syd1.au.unkin.net host has no CA bundle mounted in the jellyfin pods, so the OIDC discovery fetch fails TLS handshake (PartialChain). Authentik's discovery response is host-relative, so the browser-facing hostname must be used, not the internal one.

- Change OidEndpoint to identity.unkin.net in fafflix plugin config
- Change OidEndpoint to identity.unkin.net in cheeztv plugin config

Reviewed-on: #470
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 12:09:31 +10:00
unkin-agent 47a2ab9152 Pin jellyfin-ha image to v0.3.2 (#468)
v0.3.0 and v0.3.1 crash-looped on Postgres migration/reader bugs and were reverted. v0.3.2 fixes both and was validated end to end against production-baseline Postgres and valkey: full migration chain completes, all previously-500 endpoints return 200, RedisTranscodeSessionStore and scan-leader gating confirmed active.

- Bump jellyfin-ha image tag v0.2.0 -> v0.3.2 in cheeztv and fafflix statefulsets

Depends on a pre-sync duplicate-username check and fresh pg_dump of both databases.

Reviewed-on: #468
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 23:08:44 +10:00
unkin-agent 4748df497a puppet: install toml into the puppetserver gem path (#467)
Catalog compiles fail with `LoadError: no such file to load -- toml`: server-side functions run in the puppetserver JRuby, whose gem path is separate from the agent CRuby path this hook installs into. puppet-prod's `profiles::puppet::gems` covers both; the hook only did the agent half.

- Install toml via `puppetserver gem`, mirroring the `puppetserver_gem` resource in puppet-prod
- Note in a comment that under `set -e` a failed install takes down an already-serving compiler

Reviewed-on: #467
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 23:08:31 +10:00
unkin-agent ba14f85e51 Pin pdbmux to v0.4.0 (#466)
v0.3.0 still serves facts from cache and collapses non-4xx upstream rejections into a 502, so clients see stale facts and lose the real status.

- Pin the pdbmux image to v0.4.0

Reviewed-on: #466
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 20:36:49 +10:00
unkin-agent 8a00ddb82c Revert jellyfin-ha to v0.2.0 (#465)
v0.3.1 crash-loops both jellyfin StatefulSets deterministically on the
RatingLevels migration (concurrent Npgsql command in progress), failing
before any schema change commits. OrderedReady updates leave ordinal-1
stuck, stranding cheeztv and fafflix single-replica with no HA.

- revert cheeztv jellyfin-ha image v0.3.1 -> v0.2.0
- revert fafflix jellyfin-ha image v0.3.1 -> v0.2.0

Unblocks the stalled StatefulSet rollout.

Reviewed-on: #465
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 21:38:08 +10:00
unkin-agent d4aed39f6a pdbmux: bump image to v0.3.0 (#464)
The deployed pin sits on v0.2.0, so pdbmux still answers malformed queries with `502 all backends failed` and resolves per-certname routes by configured backend order rather than by which backend actually owns the node.

Bump the pdbmux image pin to v0.3.0:

- Replay a unanimous upstream rejection (PuppetDB's real 400 + parse message) instead of a 502.
- Resolve per-certname routes to the node's owning backend by report freshness.
- Match backend addresses case-insensitively when redacting.

Reviewed-on: #464
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 20:45:14 +10:00
unkin-agent d6a1279efe jellyfin: bump fafflix+cheeztv to v0.3.1 (#463)
v0.3.0 (#461) crash-looped existing databases on a broken Postgres
migration path; v0.3.1 restores the migration baseline, hardens guid/FK
handling, and fixes encoding.xml parsing.

- bump cheeztv jellyfin-ha image v0.2.0 -> v0.3.1
- bump fafflix jellyfin-ha image v0.2.0 -> v0.3.1

Requires manual pre-merge database verification before merge.

Reviewed-on: #463
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 20:32:41 +10:00
unkin-agent 55af4b2f16 allow catalog-diff to compile catalogs on the puppet compilers (#462)
catalog-diff compiles a host's catalog in two environments and diffs them to validate puppet-prod changes before merge, which means compiling catalogs on behalf of other nodes via POST /puppet/v4/catalog. The compilers run the image default auth.conf, where that endpoint is denied.

- add a compiler auth.conf allowing catalog-diff.main.unkin.net to POST /puppet/v4/catalog
- add a pre-default entrypoint script seeding it into conf.d, failing hard if the source is absent
- mount both onto the compiler deployment via configMapGenerator

Reviewed-on: #462
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 16:31:08 +10:00
unkin-agent 783a3db0fd revert jellyfin-ha to v0.2.0 (#461)
v0.3.0 fails an EF-Core migration on boot (NormalizedUsername column
missing), crash-looping ordinal-1 pods so the StatefulSet rolling
update stalls and cheeztv/fafflix stay single-replica. Unblocks the
stuck rollout.

- revert cheeztv statefulset image to jellyfin-ha:v0.2.0
- revert fafflix statefulset image to jellyfin-ha:v0.2.0

Reviewed-on: #461
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 15:22:20 +10:00
unkin-agent 84f09f89ff bump jellyfin-ha to v0.3.0 (#460)
jellyfin-ha v0.3.0 is the first build tracking Jellyfin 12.0 (.NET 10 runtime, jellyfin-web 12.0, LDAP plugin 24).

- bump cheeztv jellyfin-ha image tag to v0.3.0
- bump fafflix jellyfin-ha image tag to v0.3.0

Reviewed-on: #460
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 14:23:39 +10:00
unkin-agent 5b07157eeb woodpecker: raise agent workspace PVC to 20Gi (#459)
Woodpecker's k8s backend provisions a per-pipeline workspace PVC sized by
WOODPECKER_BACKEND_K8S_VOLUME_SIZE. At 10G, large builds (e.g. .NET clone +
build output) leave too little free space for tests that hard-require free
disk headroom, failing purely on disk exhaustion.

- raise WOODPECKER_BACKEND_K8S_VOLUME_SIZE from 10G to 20Gi

Reviewed-on: #459
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 12:47:07 +10:00
unkin-agent 7aec9a9021 artifactapi: restore combine-certs + PROVIDER_CA_FILES on oauth2-proxy (#458)
**Fix-forward companion to the #457 rollback. This is NOT the current outage fix — see below.**

## The actual outage

The UI is 503 because the Authentik application slug `artifactapi` **does not exist**. OIDC discovery 404s, so oauth2-proxy exits at startup, the Service has no ready endpoints, and Traefik answers `no available server`.

```
identity.unkin.net              /application/o/artifactapi/…  404
identity.k8s.syd1.au.unkin.net  /application/o/artifactapi/…  404
identity.unkin.net              /application/o/repospawner/…  200
identity.unkin.net              /application/o/argocd/…       200
```

Root cause is upstream in **terraform-authentik**: `ci/woodpecker/push/apply` on main HEAD `4e16401` **failed**. That apply has to succeed before any argocd-apps change can help. **This PR does not fix that.**

## What this PR does fix

#456 dropped the `combine-certs` initContainer and `OAUTH2_PROXY_PROVIDER_CA_FILES`, reasoning that `identity.unkin.net` serves a publicly trusted Let's Encrypt cert and so needs no internal CA. That holds for the browser redirect but not for oauth2-proxy's own back-channel discovery/token calls.

artifactapi is the **only one of six** oauth2-proxies in the estate without it:

| app | issuer host | `PROVIDER_CA_FILES` |
|---|---|---|
| arrproxy | identity.unkin.net | yes |
| logviewer | identity.unkin.net | yes |
| mediamark | identity.unkin.net | yes |
| repospawner | identity.unkin.net | yes |
| watchstate | identity.k8s… | yes |
| **artifactapi** | identity.unkin.net | **no** |

repospawner uses the **same public `identity.unkin.net` issuer** and still needs the internal bundle, which falsifies the removal reasoning. The existing comment on that initContainer states it plainly: *"The Authentik issuer is served behind the internal unkin.net CA."*

## Changes

- Add the `combine-certs` initContainer — byte-identical to repospawner's.
- Mount the combined bundle and set `OAUTH2_PROXY_PROVIDER_CA_FILES`.
- Reload the Deployment when `vault-ca-cert` rotates.

`vault-ca-cert` already exists in the `artifactapi` namespace (`api-deployment.yaml` uses it). `kustomize build apps/base/artifactapi` succeeds.

## Risk

Trust-only and strictly additive — it appends the internal CA to the system roots. Harmless if the back channel turns out to reach a publicly trusted endpoint after all. Expected to remove the *next* blocker, surfacing as x509, once the terraform-authentik apply lands.

## Sequencing

1. Fix and re-run terraform-authentik `push/apply` so the `artifactapi` application exists.
2. Merge this.
3. Confirm `/ui/` returns 200, then close #457 unmerged.

Only merge #457 instead if the UI must come back before step 1 can be done.

Reviewed-on: #458
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-07 23:05:39 +10:00
unkin-agent c98d88c197 Put the artifactapi web UI behind Authentik oauth2-proxy (#456)
The artifactapi web UI is open to anyone who can reach the host. Front it with Authentik SSO gated on akP-artifactapi-admin, while leaving the package-manager surfaces (/api/v1, /api/v2, /v2 docker registry, /terraform, /.well-known) untouched — dnf, containerd mirrors, buildah, terraform and CI publish steps cannot do a browser flow.

- Add the oauth2-proxy ConfigMap, Deployment, Service and VMPodScrape.
- Add the oauth-credentials VaultStaticSecret.
- Point the api-route /ui rule at oauth2-proxy and add a /oauth2 rule; the catch-all / rule still goes straight to the api Service on both listeners.

Requires terraform-authentik #34 applied and the Vault kv seed first.

Reviewed-on: #456
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-07 21:01:17 +10:00
unkin-agent 520da449c9 pdbmux: bump image to v0.2.0 (#455)
## Why
v0.2.0 ships the reports/events union, aggregate summing across backends, config-file support, and the removal of the primary/prefer ranking.

## How
- Pin the pdbmux Deployment image to `v0.2.0`.
- Leave `pdbmux-env` unchanged: `PDBMUX_LISTEN`, `PDBMUX_BACKENDS`, `PDBMUX_MERGE` are the only keys v0.2.0 reads from this ConfigMap, and `PDBMUX_PRIMARY`/`PDBMUX_PREFER` are already gone (#452).

Reviewed-on: #455
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 17:05:45 +10:00
unkin-agent 2fcb3d70d5 Bump artifactapi to v3.11.2 (#453)
Why: pick up v3.11.2, which moves DB migrations onto golib/pg with no behavior change.

- Bump the artifactapi API and UI image pins from v3.11.1 to v3.11.2.

Reviewed-on: #453
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 16:04:58 +10:00
unkin-agent 1172aa3e96 pdbmux: order backends new-first and drop primary/prefer (#452)
## Why
pdbmux#9 drops primary/prefer and makes configured backend order the only tie-break, so the current `old`-first list would silently reverse which PuppetDB wins.

## How
- Order `PDBMUX_BACKENDS` with `new=http://puppetdb.puppet.svc.cluster.local:8080` first and `old=http://puppetdbapi.service.consul:8080` second, URLs unchanged.
- Drop `PDBMUX_PRIMARY` and `PDBMUX_PREFER`; both already resolve to `new`, on the deployed v0.1.0 image (estate defaults) and on pdbmux main (first-backend fallback), so the rendered behaviour is unchanged today.
- Refresh the configmap and deployment comments to describe order-based precedence.

Merge this before pdbmux#9.

Reviewed-on: #452
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 16:04:40 +10:00
unkin-agent e24b19412a woodpecker: add vimpack-ci ServiceAccount (#454)
Why: the new vimpack repo's woodpecker pipelines set `serviceAccountName: vimpack-ci`, which does not exist in the woodpecker namespace yet.

- Add bare `vimpack-ci` ServiceAccount in `apps/base/woodpecker/` and register it in the kustomization.

Reviewed-on: #454
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 13:42:33 +10:00
unkin-agent b2b82e6a6c puppet: use OPENVOXSERVER_JAVA_ARGS for puppetserver JVM args (#451)
## Why

The `ghcr.io/openvoxproject/openvoxserver` image reads `OPENVOXSERVER_JAVA_ARGS` (`/etc/default/puppetserver`: `JAVA_ARGS=$OPENVOXSERVER_JAVA_ARGS`) and ships no `PUPPETSERVER_*` alias, so our heap/JMX flags have been inert since the fork switch — live masters and compilers run the image default `-Xms1024m -Xmx1024m` with no JMX.

## How

- Rename `PUPPETSERVER_JAVA_ARGS` to `OPENVOXSERVER_JAVA_ARGS` in `puppetserver-master-config`, `puppetserver-compiler-config` and `puppetserver-init-config`.
- Rename the same var on the `perms-and-dirs` init container in `deployment_puppetserver-compiler.yaml`.
- Flag values are unchanged (`-Xms1024m -Xmx3072m` plus the JMX flags); heap tuning is a separate call.

Reviewed-on: #451
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:30:22 +10:00
unkin-agent 7f15488332 Bump encapi image to v0.1.2 (#450)
Why: pick up encapi v0.1.2, which moves DB migrations onto golib/pg with no behavior change (release pipeline green).

- Bump the encapi deployment container image from v0.1.1 to v0.1.2

Reviewed-on: #450
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:23:45 +10:00
unkin-agent 8ea2925561 Set puppetserver environment timeout to 0 (#449)
## Why

An unlimited environment timeout serves stale catalogs after code changes.

## How

- Set `OPENVOXSERVER_ENVIRONMENT_TIMEOUT: "0"` in the puppetserver master and compiler ConfigMaps.

Reviewed-on: #449
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-05 11:23:04 +10:00