Compare commits

..

37 Commits

Author SHA1 Message Date
unkin-agent 757ae5b240 Add a VictoriaLogs cluster and point logs-ingest at it (#488)
The k8s log pipeline stores to ClickHouse via NATS+vector, while the VM estate ships journald to a separate puppet-managed VictoriaLogs cluster. Consolidating on VictoriaLogs in-cluster collapses the two paths, and the logs-ingest gateway has no clients yet so it can be repointed now, ahead of the puppet change.

- add VLCluster `logs` at v1.52.0 (2 vlinsert, 2 vlselect, 3 vlstorage, 180d retention, 250Gi each on cephrbd-fast-delete)
- cap vlstorage disk use at 220GiB per node so 180d stays time-based rather than disk-bound
- repoint the logs-ingest HTTPRoute at `vlinsert-logs:9481`
- add a VictoriaLogs Grafana datasource and install its plugin

Nothing is removed here; NATS, ClickHouse, vector, logarchiver and logviewer keep running until a follow-up drops them.

Reviewed-on: #488
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:33:58 +10:00
unkin-agent 02f877540c Add gocache serve Deployment with nginx stream sidecar (#496)
`go-cache-plugin serve` binds `127.0.0.1` only, so nothing outside the pod can reach it and laptops have no way to use the S3-backed Go cache without holding RGW credentials.

- Run `go-cache-plugin serve` against the `gocache` bucket, path-style, explicit region to skip the GetBucketLocation probe
- Add an nginx sidecar stream-proxying `:9090` to the loopback plugin port, `proxy_timeout 2h`
- Publish it on PureLB `198.18.200.11`, `externalTrafficPolicy: Local` so the client IP reaches the allow rules
- Restrict to workstation + pod CIDRs: GOCACHEPROG is unauthenticated and a poisoned entry runs in every consuming build

Merge only after `docker-internal/go-cache-plugin:v0.1.0` is published.

Reviewed-on: #496
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:32:32 +10:00
unkin-agent 148dac8ca2 Declare the acme.unkin.net nameservers (#495)
The zone was seeded with an apex `NS ns1.acme.unkin.net` glued to the primary pod IP. Both were later corrected by hand, so the live RRset and the ns1 address exist only in the zone journal -- a reseed republishes the pod IP and breaks DNS-01 for every `*.unkin.net` cert. Declaring them makes git the source of truth.

- declare the two published apex NS names
- declare the in-zone ns1 address, which a seed would otherwise glue to the pod IP

Matches what the zone serves today, so applying it changes no records. Requires bind-operator v0.3.0.

Reviewed-on: #495
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:10:13 +10:00
unkin-agent 53e5846c18 Roll bind-operator to v0.3.0 (#494)
v0.3.0 converges a zone apex NS onto its declared nameservers instead of leaving the seed placeholder, which publishes a primary pod IP. The CRD moves with the image because the operator reads the new `spec.nameservers` field.

- pin the bind-operator image to v0.3.0
- pull the CRDs from the v0.3.0 tag

Reviewed-on: #494
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:06:21 +10:00
unkin-agent 9fed5decc8 bump victoria-metrics-operator chart to 0.67.3 (#493)
The vm-system overlay pins victoria-metrics-operator chart 0.57.1 (operator v0.66.1), eight operator minors behind upstream, so the cluster runs without newer CRD fields and reconciler fixes.

- bump the victoria-metrics-operator helmChart version to 0.67.3 (operator v0.74.1)

Reviewed-on: #493
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:05:47 +10:00
unkin-agent 20077f1029 Move the haproxy edge behind the external Traefik (#492)
The haproxy edge holds its own DMZ VIP, a second public entry point alongside
traefik-external that must be firewalled and DNS'd separately. Traefik can
front it with TLS passthrough, leaving haproxy's certs and backends untouched.

- Add a `traefik-external` Gateway: HTTP :80 plus Passthrough TLS :443.
- TLSRoute the 12 `fe_https.map` hostnames to haproxy:443; HTTPRoute 301s :80.
- Make the Service ClusterIP on 443 only, releasing 198.18.199.1.
- Drop `fe_http`, `be_letsencrypt` and `fe_http.map`; certs are DNS-01 only.

Client IP now reads as a Traefik pod — the Gateway provider cannot emit PROXY protocol to a TLSRoute backend. `sessionAffinity` goes too (it would pin Traefik pods, not clients); SRVNAME cookies keep persistence.

Reviewed-on: #492
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 23:38:07 +10:00
unkin-agent 426a399f31 Drop stalwart mail proxying from the haproxy edge (#491)
Stalwart was only ever a test deployment. The daemon is dead on all three
backend VMs and nothing public depends on it — `unkin.net` MX points at Google —
so the edge is proxying mail to nowhere and the tcp frontends make `defaults`
emit 20 spurious HTTP-mode warnings.

- Drop the `fe_smtp`, `fe_submission`, `fe_imap` and `fe_imaps` frontends.
- Drop the five `be_stalwart_*` backends and their map entries in `fe_http.map`/`fe_https.map`.
- Drop the now-unused 25/143/587/993 Service and container ports.

`haproxy -c` on the rendered config: exit 0, 0 warnings (was 20), 0 alerts.

Reviewed-on: #491
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 21:25:49 +10:00
unkin-agent 5341253573 Add Ceph RGW bucket for the shared Go build cache (#489)
Go builds on CI and laptops each rebuild the same packages from scratch. A
GOCACHEPROG backend needs an S3 bucket plus credentials before anything can
point at it, so provision those first. The bucket lives in the woodpecker
namespace because CI is the primary consumer and reads the Secret there.

- add Bucket and ObjectStoreUser for the shared Go build cache
- use default (replicated) placement rather than the ec target, since a build
  cache is millions of small objects
- purge and drop the bucket and user on delete; the cache is disposable

Nothing consumes the bucket yet.

Reviewed-on: #489
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:56:51 +10:00
unkin-agent d48125d699 Add job and start deadlines to the g10k-code CronJob (#490)
A g10k-code job wedged in ContainerCreating on a failed CephFS mount and never reached a terminal condition, so it stayed in the CronJob active list and `concurrencyPolicy: Forbid` skipped every following minute. No Puppet code reached the estate for 6 days, and the piled-up missed slots crossed the controller 100-slot cap into `TooManyMissedTimes`. The CronJob carried no deadlines at all.

- Cap a job at `activeDeadlineSeconds: 300` on the Job spec, so a hang is failed as `DeadlineExceeded` and drops out of the active list (healthy runs take 16-18s).
- Set `startingDeadlineSeconds: 200`, bounding missed-schedule look-back to ~3 slots so the count cannot reach 100.

Reviewed-on: #490
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:55:32 +10:00
unkin-agent abf6bfae88 Drop dead X-Frame-Options rules from the haproxy edge (#487)
The 13 `X-Frame-Options DENY if acl_<host>` rules in `fe_https` have never fired:
their ACLs use `req.hdr(host)`, a request-direction fetch that is invalid in a
response ruleset, so HAProxy rejects them at config-check time. Carried over
verbatim from the Puppet LXD config during the k8s move.

- Remove the 13 dead `http-response set-header X-Frame-Options` rules.
- Remove the 15 now-orphaned `acl acl_*` definition lines.

Not switching the header on: it has never been live, and Grafana/Gitea send their
own. `haproxy -c` warnings drop 33 -> 20; the two working response headers stay.

Reviewed-on: #487
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:36:26 +10:00
unkin-agent 4190785389 Move the au-syd1 haproxy edge into Kubernetes (#485)
The au-syd1 edge proxy runs on a hand-managed LXD container outside the cluster, with no HA and no shared config source.

- Add `apps/base/haproxy/`: 3 replicas behind the DMZ LoadBalancer 198.18.199.1, config from a ConfigMap, wildcard certs from reflected secrets.
- Keep source IPs via `externalTrafficPolicy: Local`; `sessionAffinity: ClientIP` stands in for the stick-table peers a Deployment cannot name.
- Drain on shutdown: `hard-stop-after 2m`, a preStop SIGUSR1 soft-stop, 150s grace.
- Bind the stats listener to 127.0.0.1 so it is port-forward only.
- Register the app in the platform project and ApplicationSet.

Reviewed-on: #485
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 18:22:57 +10:00
unkin-agent 812a9a2f2b Publish real delegation records for acme.unkin.net (#486)
The acme.unkin.net zone still serves only the bind-operator seed apex: NS ns1.acme.unkin.net glued to A 10.42.6.38, a pod IP no pod holds. The parent delegates to acme-ns1.unkin.net, but public resolvers have already promoted the child NS RRset, so when the cached address expires DNS-01 fails for every unkin.net wildcard at once.

- Add apex NS acme-ns1.unkin.net., matching the parent delegation (out of zone, no glue needed).
- Point ns1.acme.unkin.net at 103.216.191.185 so resolvers holding the seeded NS name still reach the zone.
- The operator seed placeholder itself is tracked separately in bind-operator.

---------

Co-authored-by: unkin-agent <agent@unkin.net>
Reviewed-on: #486
Co-authored-by: Unkin Agent <unkin-agent@unkin.net>
Co-committed-by: Unkin Agent <unkin-agent@unkin.net>
2026-09-26 16:42:02 +10:00
unkin-agent f37749523d Add *.main and *.ceph wildcard certificates for haproxy (#484)
The haproxy edge terminates TLS for hosts under `main.unkin.net` and `ceph.unkin.net`, which the single `*.unkin.net` wildcard does not cover.

- Add cert-manager Certificates for both wildcards from the `letsencrypt` ClusterIssuer.
- Reflect the minted secrets into the `haproxy` namespace.

Needs these records in the public unkin.net zone first:
`_acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.`
`_acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.`

Reviewed-on: #484
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 16:38:31 +10:00
unkin-agent fe51aa07be Give the puppetserver compilers the Vault cert helpers (#482)
profiles::pki::vault and profiles::ssh::sign shell out to
/usr/local/bin/certmanager and /usr/local/bin/sshsignhost from generate()
during catalog compilation. Neither binary exists in the compiler image, so
every node using them fails to compile.

- install certmanager v0.2.0 and sshsignhost v0.1.0 onto the shared bin volume with sha256 verification
- wrap both at /usr/local/bin from a pre-default entrypoint hook, failing startup loudly if either is missing
- mount read-only Vault configs for both: kubernetes auth on k8s/au/syd1, internal CA verified rather than skipped

Reviewed-on: #482
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-24 21:15:09 +10:00
unkin-agent cdaab736b5 Bump jellyfin-ha to v0.4.0 (#483)
The deployed v0.3.3 build returns 500 from /Shows/NextUp on PostgreSQL, breaking the home screen, and lets replicas diverge: library-visibility and shared-config changes never propagate, user data (resume, played state, favourites, ratings) is overwritten between pods, and eight scheduled tasks run on every replica instead of only the scan leader. v0.4.0 carries the fixes.

- Pin cheeztv and fafflix to jellyfin-ha:v0.4.0.

No config change needed: cross-pod invalidation reuses the transcode-store Redis connection string both apps already set.

Reviewed-on: #483
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-22 22:57:06 +10:00
unkin-agent b31517e6d9 Merge pull request #481 from benvin/jellyfin-sso-valkey-state
Roll jellyfin-ha to v0.3.3 and drop Service session affinity

Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 14:02:48 +10:00
unkin-agent 5a74b2cec6 Roll bind-operator to v0.2.7 (journal-aware zone seeding) (#480)
Deploy bind-operator v0.2.7. The operator seeded a fresh skeleton zone file at serial 1 over zones whose BIND journal was still on disk at a higher serial; BIND rejected the inconsistent pair (`addzone failed: out of range`) and, with a PVC per replica, the stale journal outlived restarts while every reconcile rewrote the skeleton, so it never converged. That SERVFAILed roughly 1 in 3 authoritative answers for k8s.syd1.au.unkin.net and resolvers cached the failures.

- Bumps the operator image to v0.2.7
- Bumps the CRD install pin to the v0.2.7 tag, which changes the CRDs

Expect one rolling restart of the operator Deployment as the new image lands.

Reviewed-on: #480
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 00:53:47 +10:00
unkin-agent f14bcc4d2a Add woodpecker ServiceAccount for jellyfin-plugin-sso CI (#479)
The new `unkin/jellyfin-plugin-sso` fork is getting a Woodpecker pipeline, and its build step will set `serviceAccountName: jellyfin-plugin-sso`. Without the SA declared here the pipeline pods fail to schedule.

- add a bare ServiceAccount `jellyfin-plugin-sso` in the `woodpecker` namespace
- register it in the woodpecker base kustomization

The step only builds .NET code, so no Vault kube-auth role or RBAC is needed.

Reviewed-on: #479
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 17:32:31 +10:00
unkin-agent b01e4c3241 Rewrite slash-less Authentik token endpoint to the canonical path (#478)
Authentik advertises the token endpoint with a trailing slash, but some OIDC clients (the ArgoCD iOS app) POST to /application/o/token without one; Django's APPEND_SLASH will not redirect a POST, so the token exchange gets 405 and login fails.

- Add an exact-match rule on /application/o/token to the authentik and authentik-internal HTTPRoutes.
- Rewrite it to /application/o/token/ with a URLRewrite ReplaceFullPath filter, preserving the method and the authentik-server backend.
- Leave the catch-all PathPrefix rule untouched; exact matches outrank it in Gateway API precedence.

Reviewed-on: #478
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 16:10:30 +10:00
unkin-agent bbd5bdaa95 Enable PKCE for ArgoCD OIDC login (#477)
The Authentik client for ArgoCD is now public (the iOS app can't hold
a secret), so Authentik no longer enforces client_secret on token
exchange. PKCE replaces that as the protection against
authorization-code interception.

- Add `enablePKCEAuthentication: true` to the `oidc.config` block in
  `argocd-cm-patch.yaml`
- Note why PKCE is needed now that the client is public

Reviewed-on: #477
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 16:10:05 +10:00
unkin-agent 4762cf9e03 Add VMPodScrape for authentik-server metrics (#476)
Authentik server pods expose django_prometheus metrics on :9300, but only ldap-outpost and redis-exporter are scraped in this namespace. Add the missing per-app scrape.

- add apps/base/authentik/server-vmpodscrape.yaml selecting app.kubernetes.io/name=authentik, component=server on the metrics port
- wire it into apps/base/authentik/kustomization.yaml

Reviewed-on: #476
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 15:34:26 +10:00
unkin-agent c83a886e74 Enable pg_stat_statements on the authentik postgres cluster (#475)
The cluster preloads no statement-statistics library, so there is no per-query cost attribution in postgres and slow query paths have to be inferred from application-side metrics instead of read straight out of the database.

- preload `pg_stat_statements`
- set `pg_stat_statements.max` and `.track`, which is what makes CNPG manage the extension and create it in every database

Requires a postgres restart. Stacked on `benvin/authentik-cnpg-resources`.

Reviewed-on: #475
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:48:45 +10:00
unkin-agent 9a7200636c Raise authentik postgres CPU, memory and buffer sizing (#474)
The 500m CPU limit is a 50ms CFS quota per 100ms period, so the postgres pods are throttled on every burst even at ~0.01 cores average and each query pays that latency. 128MB of `shared_buffers` and a 256MB cache estimate also under-serve the planner on the joins authentik issues on its hot read paths.

- raise resources to requests `500m`/`1Gi`, limits `2`/`2Gi`
- raise `shared_buffers` to 512MB and `effective_cache_size` to 1536MB
- hold the post-incident memory headroom multiple over `shared_buffers`

Rolling restart with switchover. Stacked on `benvin/authentik-hot-standby-feedback`.

Reviewed-on: #474
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:35:16 +10:00
unkin-agent 9535bad9bc Enable hot_standby_feedback on the authentik postgres cluster (#473)
Authentik serves multi-second API reads from the CNPG hot standbys. Those reads outlive `max_standby_streaming_delay`, so recovery cancels them with `canceling statement due to conflict with recovery`, which authentik surfaces as HTTP 500 — enough to break a terraform apply mid-run.

- set `hot_standby_feedback` on so replicas report their oldest xmin to the primary and long reads stop being cancelled
- SIGHUP reload only, no restart or switchover
- retained-dead-tuple cost is negligible on a ~155MB database

Reviewed-on: #473
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 14:34:33 +10:00
unkin-agent 34dd70435e Auto-reload cheeztv and fafflix on plugin ConfigMap change (#471)
Edits to the cheeztv/fafflix plugin ConfigMaps only reach the pods via the inject-plugin-config initContainer, so a config change sat inert until someone manually rolled the StatefulSet. Reloader is deployed cluster-wide with autoReloadAll disabled, so each workload has to opt in.

- annotate both StatefulSets with configmap.reloader.stakater.com/auto: "true"

Reviewed-on: #471
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 12:25:08 +10:00
unkin-agent 6cc752336e Point jellyfin SSO at public Authentik hostname (#470)
The internal-CA identity.k8s.syd1.au.unkin.net host has no CA bundle mounted in the jellyfin pods, so the OIDC discovery fetch fails TLS handshake (PartialChain). Authentik's discovery response is host-relative, so the browser-facing hostname must be used, not the internal one.

- Change OidEndpoint to identity.unkin.net in fafflix plugin config
- Change OidEndpoint to identity.unkin.net in cheeztv plugin config

Reviewed-on: #470
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-19 12:09:31 +10:00
unkin-agent 47a2ab9152 Pin jellyfin-ha image to v0.3.2 (#468)
v0.3.0 and v0.3.1 crash-looped on Postgres migration/reader bugs and were reverted. v0.3.2 fixes both and was validated end to end against production-baseline Postgres and valkey: full migration chain completes, all previously-500 endpoints return 200, RedisTranscodeSessionStore and scan-leader gating confirmed active.

- Bump jellyfin-ha image tag v0.2.0 -> v0.3.2 in cheeztv and fafflix statefulsets

Depends on a pre-sync duplicate-username check and fresh pg_dump of both databases.

Reviewed-on: #468
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 23:08:44 +10:00
unkin-agent 4748df497a puppet: install toml into the puppetserver gem path (#467)
Catalog compiles fail with `LoadError: no such file to load -- toml`: server-side functions run in the puppetserver JRuby, whose gem path is separate from the agent CRuby path this hook installs into. puppet-prod's `profiles::puppet::gems` covers both; the hook only did the agent half.

- Install toml via `puppetserver gem`, mirroring the `puppetserver_gem` resource in puppet-prod
- Note in a comment that under `set -e` a failed install takes down an already-serving compiler

Reviewed-on: #467
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 23:08:31 +10:00
unkin-agent ba14f85e51 Pin pdbmux to v0.4.0 (#466)
v0.3.0 still serves facts from cache and collapses non-4xx upstream rejections into a 502, so clients see stale facts and lose the real status.

- Pin the pdbmux image to v0.4.0

Reviewed-on: #466
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-13 20:36:49 +10:00
unkin-agent 8a00ddb82c Revert jellyfin-ha to v0.2.0 (#465)
v0.3.1 crash-loops both jellyfin StatefulSets deterministically on the
RatingLevels migration (concurrent Npgsql command in progress), failing
before any schema change commits. OrderedReady updates leave ordinal-1
stuck, stranding cheeztv and fafflix single-replica with no HA.

- revert cheeztv jellyfin-ha image v0.3.1 -> v0.2.0
- revert fafflix jellyfin-ha image v0.3.1 -> v0.2.0

Unblocks the stalled StatefulSet rollout.

Reviewed-on: #465
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 21:38:08 +10:00
unkin-agent d4aed39f6a pdbmux: bump image to v0.3.0 (#464)
The deployed pin sits on v0.2.0, so pdbmux still answers malformed queries with `502 all backends failed` and resolves per-certname routes by configured backend order rather than by which backend actually owns the node.

Bump the pdbmux image pin to v0.3.0:

- Replay a unanimous upstream rejection (PuppetDB's real 400 + parse message) instead of a 502.
- Resolve per-certname routes to the node's owning backend by report freshness.
- Match backend addresses case-insensitively when redacting.

Reviewed-on: #464
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 20:45:14 +10:00
unkin-agent d6a1279efe jellyfin: bump fafflix+cheeztv to v0.3.1 (#463)
v0.3.0 (#461) crash-looped existing databases on a broken Postgres
migration path; v0.3.1 restores the migration baseline, hardens guid/FK
handling, and fixes encoding.xml parsing.

- bump cheeztv jellyfin-ha image v0.2.0 -> v0.3.1
- bump fafflix jellyfin-ha image v0.2.0 -> v0.3.1

Requires manual pre-merge database verification before merge.

Reviewed-on: #463
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 20:32:41 +10:00
unkin-agent 55af4b2f16 allow catalog-diff to compile catalogs on the puppet compilers (#462)
catalog-diff compiles a host's catalog in two environments and diffs them to validate puppet-prod changes before merge, which means compiling catalogs on behalf of other nodes via POST /puppet/v4/catalog. The compilers run the image default auth.conf, where that endpoint is denied.

- add a compiler auth.conf allowing catalog-diff.main.unkin.net to POST /puppet/v4/catalog
- add a pre-default entrypoint script seeding it into conf.d, failing hard if the source is absent
- mount both onto the compiler deployment via configMapGenerator

Reviewed-on: #462
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 16:31:08 +10:00
unkin-agent 783a3db0fd revert jellyfin-ha to v0.2.0 (#461)
v0.3.0 fails an EF-Core migration on boot (NormalizedUsername column
missing), crash-looping ordinal-1 pods so the StatefulSet rolling
update stalls and cheeztv/fafflix stay single-replica. Unblocks the
stuck rollout.

- revert cheeztv statefulset image to jellyfin-ha:v0.2.0
- revert fafflix statefulset image to jellyfin-ha:v0.2.0

Reviewed-on: #461
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 15:22:20 +10:00
unkin-agent 84f09f89ff bump jellyfin-ha to v0.3.0 (#460)
jellyfin-ha v0.3.0 is the first build tracking Jellyfin 12.0 (.NET 10 runtime, jellyfin-web 12.0, LDAP plugin 24).

- bump cheeztv jellyfin-ha image tag to v0.3.0
- bump fafflix jellyfin-ha image tag to v0.3.0

Reviewed-on: #460
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 14:23:39 +10:00
unkin-agent 5b07157eeb woodpecker: raise agent workspace PVC to 20Gi (#459)
Woodpecker's k8s backend provisions a per-pipeline workspace PVC sized by
WOODPECKER_BACKEND_K8S_VOLUME_SIZE. At 10G, large builds (e.g. .NET clone +
build output) leave too little free space for tests that hard-require free
disk headroom, failing purely on disk exhaustion.

- raise WOODPECKER_BACKEND_K8S_VOLUME_SIZE from 10G to 20Gi

Reviewed-on: #459
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-12 12:47:07 +10:00
unkin-agent 7aec9a9021 artifactapi: restore combine-certs + PROVIDER_CA_FILES on oauth2-proxy (#458)
**Fix-forward companion to the #457 rollback. This is NOT the current outage fix — see below.**

## The actual outage

The UI is 503 because the Authentik application slug `artifactapi` **does not exist**. OIDC discovery 404s, so oauth2-proxy exits at startup, the Service has no ready endpoints, and Traefik answers `no available server`.

```
identity.unkin.net              /application/o/artifactapi/…  404
identity.k8s.syd1.au.unkin.net  /application/o/artifactapi/…  404
identity.unkin.net              /application/o/repospawner/…  200
identity.unkin.net              /application/o/argocd/…       200
```

Root cause is upstream in **terraform-authentik**: `ci/woodpecker/push/apply` on main HEAD `4e16401` **failed**. That apply has to succeed before any argocd-apps change can help. **This PR does not fix that.**

## What this PR does fix

#456 dropped the `combine-certs` initContainer and `OAUTH2_PROXY_PROVIDER_CA_FILES`, reasoning that `identity.unkin.net` serves a publicly trusted Let's Encrypt cert and so needs no internal CA. That holds for the browser redirect but not for oauth2-proxy's own back-channel discovery/token calls.

artifactapi is the **only one of six** oauth2-proxies in the estate without it:

| app | issuer host | `PROVIDER_CA_FILES` |
|---|---|---|
| arrproxy | identity.unkin.net | yes |
| logviewer | identity.unkin.net | yes |
| mediamark | identity.unkin.net | yes |
| repospawner | identity.unkin.net | yes |
| watchstate | identity.k8s… | yes |
| **artifactapi** | identity.unkin.net | **no** |

repospawner uses the **same public `identity.unkin.net` issuer** and still needs the internal bundle, which falsifies the removal reasoning. The existing comment on that initContainer states it plainly: *"The Authentik issuer is served behind the internal unkin.net CA."*

## Changes

- Add the `combine-certs` initContainer — byte-identical to repospawner's.
- Mount the combined bundle and set `OAUTH2_PROXY_PROVIDER_CA_FILES`.
- Reload the Deployment when `vault-ca-cert` rotates.

`vault-ca-cert` already exists in the `artifactapi` namespace (`api-deployment.yaml` uses it). `kustomize build apps/base/artifactapi` succeeds.

## Risk

Trust-only and strictly additive — it appends the internal CA to the system roots. Harmless if the back channel turns out to reach a publicly trusted endpoint after all. Expected to remove the *next* blocker, surfacing as x509, once the terraform-authentik apply lands.

## Sequencing

1. Fix and re-run terraform-authentik `push/apply` so the `artifactapi` application exists.
2. Merge this.
3. Confirm `/ui/` returns 200, then close #457 unmerged.

Only merge #457 instead if the UI must come back before step 1 can be done.

Reviewed-on: #458
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-07 23:05:39 +10:00
64 changed files with 1860 additions and 32 deletions
+28 -1
View File
@@ -1,4 +1,21 @@
--- ---
# Path split between the authenticated UI and the unauthenticated machine API.
# Longest matching prefix wins, so the two UI rules take precedence over "/".
#
# AUTHENTICATED (oauth2 Service -> oauth2-proxy -> ui Service):
# /oauth2 oauth2-proxy sign_in / start / callback / sign_out
# /ui the human-facing SPA
#
# NOT AUTHENTICATED (artifactapi Service, unchanged):
# /api/v1/{remote,local,virtual}/* package proxy reads (yum/dnf, pip, ...)
# /api/v2/remotes|virtuals|locals/* management API + the UI's own XHR calls
# /api/v2/remotes/{name}/files/* CI publish uploads (PUT) and downloads
# /v2/* Docker Registry V2 (containerd, buildah)
# /terraform/v1/providers/* Terraform provider registry
# /.well-known/terraform.json Terraform service discovery
# /health, /version, / probes and the redirect to /ui/
# Those clients cannot complete a browser OIDC flow, so they must never be
# routed through oauth2-proxy.
apiVersion: gateway.networking.k8s.io/v1 apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute kind: HTTPRoute
metadata: metadata:
@@ -22,7 +39,17 @@ spec:
- backendRefs: - backendRefs:
- group: "" - group: ""
kind: Service kind: Service
name: ui name: oauth2
port: 80
weight: 1
matches:
- path:
type: PathPrefix
value: /oauth2
- backendRefs:
- group: ""
kind: Service
name: oauth2
port: 80 port: 80
weight: 1 weight: 1
matches: matches:
+2
View File
@@ -12,6 +12,8 @@ resources:
- gateway.yaml - gateway.yaml
- httproute.yaml - httproute.yaml
- namespace.yaml - namespace.yaml
- oauth2-proxy-configmap.yaml
- oauth2-proxy-deployment.yaml
- redis-deployment.yaml - redis-deployment.yaml
- services.yaml - services.yaml
- ui-deployment.yaml - ui-deployment.yaml
@@ -0,0 +1,46 @@
---
# Non-secret oauth2-proxy configuration (client_id/secret/cookie_secret come
# from the oauth-credentials Secret).
#
# SCOPE: this proxy fronts the artifactapi web UI ONLY. The HTTPRoute sends just
# /ui and /oauth2 here; every machine surface (/api/v1, /api/v2, /v2 docker
# registry, /terraform, /.well-known/terraform.json, /health, /version, /) goes
# straight to the api Service and is NOT authenticated. yum/dnf, containerd
# registry mirrors, docker/buildah, terraform init and Woodpecker publish steps
# cannot complete a browser OIDC flow, so they must never reach this container.
# Its only upstream is the ui Service -- there is deliberately no api upstream.
apiVersion: v1
kind: ConfigMap
metadata:
name: artifactapi-oauth2-env
namespace: artifactapi
data:
OAUTH2_PROXY_HTTP_ADDRESS: "0.0.0.0:4180"
OAUTH2_PROXY_METRICS_ADDRESS: "0.0.0.0:44180"
OAUTH2_PROXY_PROVIDER: "oidc"
# Publicly-trusted Authentik host: the authorize step is a browser redirect,
# so the issuer must present a cert every user's browser already trusts (the
# k8s host serves an internal-CA cert). Slug from terraform-authentik.
OAUTH2_PROXY_OIDC_ISSUER_URL: "https://identity.unkin.net/application/o/artifactapi/"
OAUTH2_PROXY_REDIRECT_URL: "https://artifactapi.k8s.syd1.au.unkin.net/oauth2/callback"
OAUTH2_PROXY_UPSTREAMS: "http://ui.artifactapi.svc.cluster.local:80/"
OAUTH2_PROXY_SCOPE: "openid email profile ak_groups"
# Populate session.Groups from the Authentik hierarchical ak_groups claim.
OAUTH2_PROXY_OIDC_GROUPS_CLAIM: "ak_groups"
OAUTH2_PROXY_ALLOWED_GROUPS: "akP-artifactapi-admin"
OAUTH2_PROXY_PASS_USER_HEADERS: "true"
OAUTH2_PROXY_EMAIL_DOMAINS: "*"
# Authentik hardcodes email_verified=false in the id_token; authorization is
# enforced via ak_groups, so accepting the unverified email is safe.
OAUTH2_PROXY_INSECURE_OIDC_ALLOW_UNVERIFIED_EMAIL: "true"
OAUTH2_PROXY_COOKIE_SECURE: "true"
OAUTH2_PROXY_COOKIE_DOMAINS: "artifactapi.k8s.syd1.au.unkin.net"
OAUTH2_PROXY_WHITELIST_DOMAINS: "artifactapi.k8s.syd1.au.unkin.net"
OAUTH2_PROXY_REVERSE_PROXY: "true"
OAUTH2_PROXY_CODE_CHALLENGE_METHOD: "S256"
OAUTH2_PROXY_SKIP_PROVIDER_BUTTON: "true"
# Back-channel discovery/token calls resolve the issuer inside the cluster,
# where it is served under the internal unkin.net CA rather than the publicly
# trusted cert the browser sees. Trust the bundle the combine-certs init
# container assembles, as every other oauth2-proxy in the estate does.
OAUTH2_PROXY_PROVIDER_CA_FILES: "/etc/ssl/combined/ca-certificates.crt"
@@ -0,0 +1,136 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: oauth2
namespace: artifactapi
annotations:
configmap.reloader.stakater.com/auto: "true"
secret.reloader.stakater.com/reload: "oauth-credentials,vault-ca-cert"
spec:
replicas: 2
selector:
matchLabels:
app: oauth2
strategy:
rollingUpdate:
maxUnavailable: 1
type: RollingUpdate
template:
metadata:
labels:
app: oauth2
spec:
serviceAccountName: default
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 65532
runAsGroup: 65532
fsGroup: 65532
seccompProfile:
type: RuntimeDefault
initContainers:
# The Authentik issuer is served behind the internal unkin.net CA;
# combine the system roots with it so oauth2-proxy's OIDC HTTP client
# trusts the discovery endpoint.
- name: combine-certs
image: docker.io/library/alpine:3
imagePullPolicy: IfNotPresent
command:
- sh
- -c
- cat /etc/ssl/certs/ca-certificates.crt /custom-ca/ca.crt > /combined-certs/ca-certificates.crt
volumeMounts:
- name: vault-ca-cert
mountPath: /custom-ca
readOnly: true
- name: combined-certs
mountPath: /combined-certs
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 50m
memory: 32Mi
limits:
cpu: 200m
memory: 64Mi
containers:
- name: oauth2-proxy
image: quay.io/oauth2-proxy/oauth2-proxy:v7.15.3
imagePullPolicy: IfNotPresent
ports:
- containerPort: 4180
name: http
protocol: TCP
- containerPort: 44180
name: metrics
protocol: TCP
envFrom:
- configMapRef:
name: artifactapi-oauth2-env
optional: false
env:
- name: OAUTH2_PROXY_CLIENT_ID
valueFrom:
secretKeyRef:
name: oauth-credentials
key: client_id
- name: OAUTH2_PROXY_CLIENT_SECRET
valueFrom:
secretKeyRef:
name: oauth-credentials
key: client_secret
- name: OAUTH2_PROXY_COOKIE_SECRET
valueFrom:
secretKeyRef:
name: oauth-credentials
key: cookie_secret
livenessProbe:
httpGet:
path: /ping
port: http
initialDelaySeconds: 10
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: http
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
volumeMounts:
- name: combined-certs
mountPath: /etc/ssl/combined
readOnly: true
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
volumes:
- name: vault-ca-cert
secret:
secretName: vault-ca-cert
items:
- key: ca.crt
path: ca.crt
- name: combined-certs
emptyDir: {}
restartPolicy: Always
+20
View File
@@ -16,6 +16,26 @@ spec:
sessionAffinity: None sessionAffinity: None
type: ClusterIP type: ClusterIP
--- ---
# Authenticated front door for the web UI only: api-route sends /ui and /oauth2
# here, oauth2-proxy authenticates and forwards to the ui Service. Every other
# path reaches the api Service above directly and stays unauthenticated.
apiVersion: v1
kind: Service
metadata:
name: oauth2
namespace: artifactapi
spec:
internalTrafficPolicy: Cluster
ports:
- name: http
port: 80
protocol: TCP
targetPort: http
selector:
app: oauth2
sessionAffinity: None
type: ClusterIP
---
apiVersion: v1 apiVersion: v1
kind: Service kind: Service
metadata: metadata:
@@ -32,3 +32,26 @@ spec:
refreshAfter: 5m refreshAfter: 5m
type: kv-v2 type: kv-v2
vaultAuthRef: default vaultAuthRef: default
---
# Authentik OIDC client for the artifactapi UI front door (client_id,
# client_secret, cookie_secret). Seeded out of band at
# kv/kubernetes/namespace/artifactapi/default/oauth-credentials; the default
# k8s auth role already grants the artifactapi/default ServiceAccount read on
# kv/data/kubernetes/namespace/{{sa_namespace}}/{{sa_name}}/*, so no
# terraform-vault change is needed. Consumed by the oauth2 Deployment.
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultStaticSecret
metadata:
name: oauth-credentials
namespace: artifactapi
spec:
destination:
create: true
name: oauth-credentials
overwrite: true
hmacSecretData: true
mount: kv
path: kubernetes/namespace/artifactapi/default/oauth-credentials
refreshAfter: 5m
type: kv-v2
vaultAuthRef: default
+14
View File
@@ -14,3 +14,17 @@ spec:
podMetricsEndpoints: podMetricsEndpoints:
- port: metrics - port: metrics
path: /metrics path: /metrics
---
# Scrape the UI oauth2-proxy (:44180), which exposes sign-in/authz counters.
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: oauth2
namespace: artifactapi
spec:
selector:
matchLabels:
app: oauth2
podMetricsEndpoints:
- port: metrics
path: /metrics
+21 -6
View File
@@ -64,8 +64,12 @@ spec:
archive_mode: "on" archive_mode: "on"
archive_timeout: 5min archive_timeout: 5min
dynamic_shared_memory_type: posix dynamic_shared_memory_type: posix
effective_cache_size: 256MB effective_cache_size: 1536MB
full_page_writes: "on" full_page_writes: "on"
# Replicas report their oldest xmin to the primary, so multi-second reads on
# a hot standby stop exhausting max_standby_streaming_delay and being
# cancelled. Retained-dead-tuple cost is negligible on a ~155MB database.
hot_standby_feedback: "on"
log_destination: csvlog log_destination: csvlog
log_directory: /controller/log log_directory: /controller/log
log_filename: postgres log_filename: postgres
@@ -77,7 +81,12 @@ spec:
max_parallel_workers: "16" max_parallel_workers: "16"
max_replication_slots: "16" max_replication_slots: "16"
max_worker_processes: "16" max_worker_processes: "16"
shared_buffers: 128MB # A pg_stat_statements.* parameter is what makes CNPG treat the extension as
# managed and run CREATE EXTENSION in every database; preloading alone does
# not create it.
pg_stat_statements.max: "10000"
pg_stat_statements.track: top
shared_buffers: 512MB
shared_memory_type: mmap shared_memory_type: mmap
ssl_max_protocol_version: TLSv1.3 ssl_max_protocol_version: TLSv1.3
ssl_min_protocol_version: TLSv1.3 ssl_min_protocol_version: TLSv1.3
@@ -86,6 +95,9 @@ spec:
wal_log_hints: "on" wal_log_hints: "on"
wal_receiver_timeout: 5s wal_receiver_timeout: 5s
wal_sender_timeout: 5s wal_sender_timeout: 5s
# CNPG merges this with the libraries it manages itself.
shared_preload_libraries:
- pg_stat_statements
syncReplicaElectionConstraint: syncReplicaElectionConstraint:
enabled: false enabled: false
primaryUpdateMethod: restart primaryUpdateMethod: restart
@@ -105,13 +117,16 @@ spec:
updateInterval: 30 updateInterval: 30
resources: resources:
limits: limits:
cpu: 500m # 500m is a 50ms CFS quota per 100ms period, exhausted by bursts even at
# ~0.01 cores average, so every query pays throttle latency.
cpu: "2"
# 512Mi OOMKilled replicas under load (shared_buffers 128MB + # 512Mi OOMKilled replicas under load (shared_buffers 128MB +
# max_connections 200 leave no headroom) — see incident 2026-07-28. # max_connections 200 leave no headroom) — see incident 2026-07-28.
memory: 1Gi # shared_buffers 512MB needs the same headroom multiple, hence 2Gi.
memory: 2Gi
requests: requests:
cpu: 50m cpu: 500m
memory: 512Mi memory: 1Gi
smartShutdownTimeout: 180 smartShutdownTimeout: 180
startDelay: 3600 startDelay: 3600
stopDelay: 1800 stopDelay: 1800
+32
View File
@@ -37,6 +37,22 @@ spec:
name: authentik name: authentik
sectionName: https sectionName: https
rules: rules:
- backendRefs:
- group: ""
kind: Service
name: authentik-server
port: 80
weight: 1
filters:
- type: URLRewrite
urlRewrite:
path:
type: ReplaceFullPath
replaceFullPath: /application/o/token/
matches:
- path:
type: Exact
value: /application/o/token
- backendRefs: - backendRefs:
- group: "" - group: ""
kind: Service kind: Service
@@ -86,6 +102,22 @@ spec:
name: authentik-internal name: authentik-internal
sectionName: https sectionName: https
rules: rules:
- backendRefs:
- group: ""
kind: Service
name: authentik-server
port: 80
weight: 1
filters:
- type: URLRewrite
urlRewrite:
path:
type: ReplaceFullPath
replaceFullPath: /application/o/token/
matches:
- path:
type: Exact
value: /application/o/token
- backendRefs: - backendRefs:
- group: "" - group: ""
kind: Service kind: Service
+1
View File
@@ -19,6 +19,7 @@ resources:
- redis-deployment.yaml - redis-deployment.yaml
- redis-pvc.yaml - redis-pvc.yaml
- redis-service.yaml - redis-service.yaml
- server-vmpodscrape.yaml
- vaultauth.yaml - vaultauth.yaml
- vaultstaticsecret.yaml - vaultstaticsecret.yaml
- vmpodscrape.yaml - vmpodscrape.yaml
@@ -0,0 +1,16 @@
---
# Scrape the authentik server's django_prometheus endpoint (:9300). Picked up
# by the observability VMAgent (selectAllByDefault).
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: authentik-server
namespace: authentik
spec:
selector:
matchLabels:
app.kubernetes.io/name: authentik
app.kubernetes.io/component: server
podMetricsEndpoints:
- port: metrics
path: /metrics
@@ -7,4 +7,5 @@ resources:
- cluster.yaml - cluster.yaml
- tsigkey.yaml - tsigkey.yaml
- zones.yaml - zones.yaml
- records.yaml
- agent-dns-rolebinding.yaml - agent-dns-rolebinding.yaml
+36
View File
@@ -0,0 +1,36 @@
# Authoritative delegation records for acme.unkin.net. Without these the zone
# only holds the operator's seed apex (NS ns1.acme.unkin.net glued to the
# primary pod IP), which is unroutable off-cluster and goes stale on
# reschedule. DNSRecords must live in the same namespace as their BindZone.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: acme-apex-ns
namespace: bind-external
spec:
zoneRef: acme-unkin-net
# "@" is the zone apex.
name: "@"
type: NS
ttl: 3600
values:
# Matches the parent delegation in Google Cloud DNS. Out of zone, so the
# child needs no glue of its own.
- acme-ns1.unkin.net.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: acme-ns1-a
namespace: bind-external
spec:
zoneRef: acme-unkin-net
name: ns1
type: A
ttl: 3600
values:
# Public address of this cluster's external BIND, same target as
# acme-ns1.unkin.net. Resolvers that cached the seeded ns1.acme.unkin.net
# NS name must still reach the zone.
- 103.216.191.185
+11
View File
@@ -17,3 +17,14 @@ spec:
updateKeyRef: certmanager updateKeyRef: certmanager
allowTransfer: allowTransfer:
- key certmanager - key certmanager
# Published apex NS. acme-ns1 is what the parent delegates to and glues; ns1 is
# in-zone, so its address is declared below or a reseed would glue it to the
# primary pod IP.
nameservers:
- acme-ns1.unkin.net.
- ns1.acme.unkin.net.
records:
- name: ns1
type: A
ttl: 3600
values: ["103.216.191.185"]
+1 -1
View File
@@ -21,7 +21,7 @@ spec:
runAsNonRoot: true runAsNonRoot: true
containers: containers:
- name: operator - name: operator
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/bind-operator:v0.2.6 image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/bind-operator:v0.3.0
args: args:
- --metrics-bind-address=:8080 - --metrics-bind-address=:8080
- --health-probe-bind-address=:8081 - --health-probe-bind-address=:8081
+1 -1
View File
@@ -6,7 +6,7 @@ resources:
- namespace.yaml - namespace.yaml
# CRDs are pulled from the bind-operator repo at the matching tag rather than # CRDs are pulled from the bind-operator repo at the matching tag rather than
# vendored here, so they never drift from the operator. # vendored here, so they never drift from the operator.
- https://git.unkin.net/unkin/bind-operator/raw/tag/v0.2.6/config/crd/install.yaml - https://git.unkin.net/unkin/bind-operator/raw/tag/v0.3.0/config/crd/install.yaml
- rbac.yaml - rbac.yaml
- agent-dns-rbac.yaml - agent-dns-rbac.yaml
- deployment.yaml - deployment.yaml
@@ -0,0 +1,26 @@
---
# Let's Encrypt *.ceph.unkin.net wildcard for the haproxy edge (ceph dashboard).
# DNS-01 needs the delegated _acme-challenge.ceph.unkin.net CNAME in the public
# unkin.net zone.
# _acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: wildcard-ceph-unkin-net
namespace: cert-manager
spec:
secretName: wildcard-ceph-unkin-net-tls
secretTemplate:
annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "haproxy"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "haproxy"
privateKey:
size: 4096
dnsNames:
- "*.ceph.unkin.net"
issuerRef:
name: letsencrypt
kind: ClusterIssuer
group: cert-manager.io
@@ -0,0 +1,26 @@
---
# Let's Encrypt *.main.unkin.net wildcard for the haproxy edge (pve, arr stack,
# jellyfin, stalwart webadmin/autoconfig). DNS-01 needs the delegated
# _acme-challenge.main.unkin.net CNAME in the public unkin.net zone.
# _acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: wildcard-main-unkin-net
namespace: cert-manager
spec:
secretName: wildcard-main-unkin-net-tls
secretTemplate:
annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "haproxy"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "haproxy"
privateKey:
size: 4096
dnsNames:
- "*.main.unkin.net"
issuerRef:
name: letsencrypt
kind: ClusterIssuer
group: cert-manager.io
@@ -14,9 +14,9 @@ spec:
secretTemplate: secretTemplate:
annotations: annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true" reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner" reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner,haproxy"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true" reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner" reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner,haproxy"
privateKey: privateKey:
size: 4096 size: 4096
dnsNames: dnsNames:
@@ -12,3 +12,5 @@ resources:
- clusterissuer_letsencrypt.yaml - clusterissuer_letsencrypt.yaml
- clusterissuer_letsencrypt-staging.yaml - clusterissuer_letsencrypt-staging.yaml
- certificate_wildcard-unkin-net.yaml - certificate_wildcard-unkin-net.yaml
- certificate_wildcard-main-unkin-net.yaml
- certificate_wildcard-ceph-unkin-net.yaml
+1 -1
View File
@@ -26,7 +26,7 @@ data:
</key> </key>
<value> <value>
<PluginConfiguration> <PluginConfiguration>
<OidEndpoint>https://identity.k8s.syd1.au.unkin.net/application/o/jellyfin/</OidEndpoint> <OidEndpoint>https://identity.unkin.net/application/o/jellyfin/</OidEndpoint>
<OidClientId>jellyfin</OidClientId> <OidClientId>jellyfin</OidClientId>
<OidSecret>@@CLIENT_SECRET@@</OidSecret> <OidSecret>@@CLIENT_SECRET@@</OidSecret>
<Enabled>true</Enabled> <Enabled>true</Enabled>
-2
View File
@@ -13,6 +13,4 @@ spec:
targetPort: http targetPort: http
selector: selector:
app: cheeztv app: cheeztv
# Pin each client to one replica to reduce transcode-session churn/takeover.
sessionAffinity: ClientIP
type: ClusterIP type: ClusterIP
+3 -1
View File
@@ -4,6 +4,8 @@ kind: StatefulSet
metadata: metadata:
name: cheeztv name: cheeztv
namespace: cheeztv namespace: cheeztv
annotations:
configmap.reloader.stakater.com/auto: "true"
spec: spec:
# HA: two replicas coordinate transcode session ownership through Valkey and # HA: two replicas coordinate transcode session ownership through Valkey and
# resume each other's HLS segments off the shared RWX transcode PVC. Stable # resume each other's HLS segments off the shared RWX transcode PVC. Stable
@@ -162,7 +164,7 @@ spec:
readOnly: true readOnly: true
containers: containers:
- name: cheeztv - name: cheeztv
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.2.0 image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.4.0
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
ports: ports:
- name: http - name: http
+1 -1
View File
@@ -26,7 +26,7 @@ data:
</key> </key>
<value> <value>
<PluginConfiguration> <PluginConfiguration>
<OidEndpoint>https://identity.k8s.syd1.au.unkin.net/application/o/jellyfin/</OidEndpoint> <OidEndpoint>https://identity.unkin.net/application/o/jellyfin/</OidEndpoint>
<OidClientId>jellyfin</OidClientId> <OidClientId>jellyfin</OidClientId>
<OidSecret>@@CLIENT_SECRET@@</OidSecret> <OidSecret>@@CLIENT_SECRET@@</OidSecret>
<Enabled>true</Enabled> <Enabled>true</Enabled>
-2
View File
@@ -13,6 +13,4 @@ spec:
targetPort: http targetPort: http
selector: selector:
app: fafflix app: fafflix
# Pin each client to one replica to reduce transcode-session churn/takeover.
sessionAffinity: ClientIP
type: ClusterIP type: ClusterIP
+3 -1
View File
@@ -4,6 +4,8 @@ kind: StatefulSet
metadata: metadata:
name: fafflix name: fafflix
namespace: fafflix namespace: fafflix
annotations:
configmap.reloader.stakater.com/auto: "true"
spec: spec:
# HA: two replicas coordinate transcode session ownership through Valkey and # HA: two replicas coordinate transcode session ownership through Valkey and
# resume each other's HLS segments off the shared RWX transcode PVC. Stable # resume each other's HLS segments off the shared RWX transcode PVC. Stable
@@ -162,7 +164,7 @@ spec:
readOnly: true readOnly: true
containers: containers:
- name: fafflix - name: fafflix
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.2.0 image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.4.0
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
ports: ports:
- name: http - name: http
+19
View File
@@ -20,3 +20,22 @@ spec:
jsonData: jsonData:
timeInterval: "15s" timeInterval: "15s"
httpMethod: "POST" httpMethod: "POST"
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDatasource
metadata:
name: victorialogs
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
plugins:
- name: victoriametrics-logs-datasource
version: 0.32.0
datasource:
name: "VictoriaLogs"
type: "victoriametrics-logs-datasource"
uid: "victorialogs"
access: "proxy"
url: "http://vlselect-logs.logging.svc.cluster.local:9471"
+274
View File
@@ -0,0 +1,274 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: haproxy-config
namespace: haproxy
data:
certificate.list: |
# First entry is the default cert for non-matching SNI.
/etc/haproxy/certs/unkin-net/tls.crt
/etc/haproxy/certs/main-unkin-net/tls.crt
/etc/haproxy/certs/ceph-unkin-net/tls.crt
fe_https.map: |
sonarr.main.unkin.net be_sonarr
radarr.main.unkin.net be_radarr
lidarr.main.unkin.net be_lidarr
readarr.main.unkin.net be_readarr
prowlarr.main.unkin.net be_prowlarr
nzbget.main.unkin.net be_nzbget
jellyfin.main.unkin.net be_jellyfin
fafflix.unkin.net be_jellyfin
git.unkin.net be_gitea
grafana.unkin.net be_grafana
dashboard.ceph.unkin.net be_ceph_dashboard
auth.unkin.net be_k8s_kanidm
haproxy.cfg: |
global
log stdout format raw local0
log stdout format raw local1 notice
maxconn 4000
hard-stop-after 2m
ssl-default-bind-ciphers EECDH+AESGCM:EDH+AESGCM:AES256+EECDH:AES256+EDH
ssl-default-bind-options ssl-min-ver TLSv1.2 ssl-max-ver TLSv1.3
ssl-default-server-ciphers kEECDH+aRSA+AES:kRSA+AES:+AES256:RC4-SHA:!kEDH:!LOW:!EXP:!MD5:!aNULL:!eNULL
ssl-default-server-options no-sslv3
stats timeout 30s
stats socket /var/lib/haproxy/stats
stats socket /var/lib/haproxy/admin.sock mode 660 level admin
tune.ssl.default-dh-param 2048
defaults
log global
maxconn 5000
mode http
option httplog
option dontlognull
option http-server-close
option forwardfor except 127.0.0.0/8
option redispatch
retries 3
stats enable
timeout http-request 10s
timeout queue 1m
timeout connect 10s
timeout client 5m
timeout server 5m
timeout http-keep-alive 10s
timeout check 10s
frontend fe_https
bind 0.0.0.0:443 ssl crt-list /usr/local/etc/haproxy/certificate.list ciphers EECDH+AESGCM:EDH+AESGCM:AES256+EECDH:AES256+EDH force-tlsv12
mode http
description Global HTTPS Frontend
http-request set-header X-Forwarded-Proto https
http-request set-header X-Real-IP %[src]
http-response set-header X-Content-Type-Options nosniff
http-response set-header X-XSS-Protection 1;mode=block
use_backend %[req.hdr(host),lower,map(/usr/local/etc/haproxy/fe_https.map,be_default)]
frontend fe_metrics
bind 0.0.0.0:8405
mode http
description Metrics Frontend
http-request set-header X-Forwarded-Proto https
http-request set-header X-Real-IP %[src]
http-request use-service prometheus-exporter if { path /metrics }
backend be_ceph_dashboard
description Backend for Ceph Dashboard from Mgr instances
balance roundrobin
cookie SRVNAME insert indirect nocache
http-check expect status 200
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 9443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
stick-table type ip size 200k expire 30m
server prodnxsr0009 198.18.23.9:9443 check cookie prodnxsr0009 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0010 198.18.23.10:9443 check cookie prodnxsr0010 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0011 198.18.23.11:9443 check cookie prodnxsr0011 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0012 198.18.23.12:9443 check cookie prodnxsr0012 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0013 198.18.23.13:9443 check cookie prodnxsr0013 fall 2 inter 2s rise 3 ssl verify none
backend be_default
description Backend for unmatched HTTP traffic
balance roundrobin
cookie SRVNAME insert
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
option httpchk GET /
option forwardfor
backend be_gitea
description Backend for gitea cluster
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
stick on src
stick-table type ip size 200k expire 30m
server ausyd1nxvm2080 198.18.26.18:443 check cookie ausyd1nxvm2080 fall 2 inter 2s rise 3 ssl verify none
server ausyd1nxvm2081 198.18.27.117:443 check cookie ausyd1nxvm2081 fall 2 inter 2s rise 3 ssl verify none
server ausyd1nxvm2082 198.18.28.71:443 check cookie ausyd1nxvm2082 fall 2 inter 2s rise 3 ssl verify none
backend be_grafana
description Backend for grafana nodes
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
stick on src
stick-table type ip size 200k expire 30m
server ausyd1nxvm2015 198.18.27.2:443 check cookie ausyd1nxvm2015 fall 2 inter 2s rise 3 ssl verify none
server ausyd1nxvm2016 198.18.28.189:443 check cookie ausyd1nxvm2016 fall 2 inter 2s rise 3 ssl verify none
backend be_jellyfin
description Backend for au-syd1 jellyfin
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2051 198.18.25.164:443 check cookie ausyd1nxvm2051 fall 2 inter 2s rise 3 ssl verify none
backend be_k8s_kanidm
description Backend for Kanidm (auth.unkin.net via Kubernetes internal Traefik)
balance roundrobin
http-reuse always
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
redirect scheme https if !{ ssl_fc }
option httpchk
option forwardfor
option http-keep-alive
option prefer-last-server
http-check connect ssl sni auth.unkin.net
http-check send meth GET uri /status ver HTTP/1.1 hdr Host auth.unkin.net
http-check expect status 200
server k8s-traefik-internal 198.18.200.4:443 ssl verify none check inter 2s rise 3 fall 2 sni str(auth.unkin.net)
backend be_lidarr
description Backend for au-syd1 lidarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2048 198.18.28.165:443 check cookie ausyd1nxvm2048 fall 2 inter 2s rise 3 ssl verify none
backend be_nzbget
description Backend for au-syd1 nzbget
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2045 198.18.25.44:443 check cookie ausyd1nxvm2045 fall 2 inter 2s rise 3 ssl verify none
backend be_prowlarr
description Backend for au-syd1 prowlarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2050 198.18.25.66:443 check cookie ausyd1nxvm2050 fall 2 inter 2s rise 3 ssl verify none
backend be_radarr
description Backend for au-syd1 radarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2047 198.18.27.131:443 check cookie ausyd1nxvm2047 fall 2 inter 2s rise 3 ssl verify none
backend be_readarr
description Backend for au-syd1 readarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2049 198.18.29.32:443 check cookie ausyd1nxvm2049 fall 2 inter 2s rise 3 ssl verify none
backend be_sonarr
description Backend for au-syd1 sonarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2046 198.18.26.161:443 check cookie ausyd1nxvm2046 fall 2 inter 2s rise 3 ssl verify none
# The `peers au-syd1-prod` section is dropped: peer names must be static and a
# Deployment cannot provide them. Behind the external Traefik's TLS
# passthrough `src` is a Traefik pod, so X-Real-IP, forwardfor and the
# `stick on src` tables all key on that; the SRVNAME cookie carries real
# session persistence. Traefik cannot emit PROXY protocol to a TLSRoute
# backend, so there is nothing to bind `accept-proxy` to.
listen health
bind 0.0.0.0:8404
mode http
monitor-uri /healthz
listen stats
bind 127.0.0.1:9090
mode http
stats uri /
stats auth admin:admin
+148
View File
@@ -0,0 +1,148 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: haproxy
namespace: haproxy
annotations:
reloader.stakater.com/auto: "true"
spec:
replicas: 3
selector:
matchLabels:
app: haproxy
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app: haproxy
spec:
automountServiceAccountToken: false
terminationGracePeriodSeconds: 150
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
app: haproxy
topologyKey: kubernetes.io/hostname
securityContext:
runAsNonRoot: true
runAsUser: 99
runAsGroup: 99
seccompProfile:
type: RuntimeDefault
containers:
- name: haproxy
image: haproxy:3.2.24-alpine
imagePullPolicy: IfNotPresent
command:
- haproxy
- -W
- -db
- -f
- /usr/local/etc/haproxy/haproxy.cfg
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: [ALL]
# fe_https binds the privileged port 443 as uid 99, and the
# dst_port ACLs need the real port.
add: [NET_BIND_SERVICE]
ports:
- name: https
containerPort: 443
protocol: TCP
- name: health
containerPort: 8404
protocol: TCP
- name: metrics
containerPort: 8405
protocol: TCP
- name: stats
containerPort: 9090
protocol: TCP
lifecycle:
preStop:
exec:
# SIGUSR1 to the master soft-stops the workers; hard-stop-after
# caps the drain. Wait so kubelet holds SIGTERM until it is done.
command:
- /bin/sh
- -c
- kill -s USR1 1; while kill -0 1 2>/dev/null; do sleep 1; done
livenessProbe:
httpGet:
path: /healthz
port: health
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /healthz
port: health
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 5
failureThreshold: 3
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: 2
memory: 1Gi
volumeMounts:
- name: config
mountPath: /usr/local/etc/haproxy
readOnly: true
- name: cert-unkin-net
mountPath: /etc/haproxy/certs/unkin-net
readOnly: true
- name: cert-main-unkin-net
mountPath: /etc/haproxy/certs/main-unkin-net
readOnly: true
- name: cert-ceph-unkin-net
mountPath: /etc/haproxy/certs/ceph-unkin-net
readOnly: true
- name: run
mountPath: /var/lib/haproxy
volumes:
- name: config
configMap:
name: haproxy-config
# ssl-load-extra-files loads <crtfile>.key by default, so the key is
# projected next to the cert as tls.crt.key.
- name: cert-unkin-net
secret:
secretName: wildcard-unkin-net-tls
items:
- key: tls.crt
path: tls.crt
- key: tls.key
path: tls.crt.key
- name: cert-main-unkin-net
secret:
secretName: wildcard-main-unkin-net-tls
items:
- key: tls.crt
path: tls.crt
- key: tls.key
path: tls.crt.key
- name: cert-ceph-unkin-net
secret:
secretName: wildcard-ceph-unkin-net-tls
items:
- key: tls.crt
path: tls.crt
- key: tls.key
path: tls.crt.key
- name: run
emptyDir: {}
restartPolicy: Always
+31
View File
@@ -0,0 +1,31 @@
---
# External (DMZ) front for the haproxy edge on the traefik-external LB VIP
# 198.18.199.0. The :443 listener is TLS Passthrough: haproxy owns the three
# wildcard certs and terminates behind Traefik, so there are no certificateRefs
# here. Listener hostnames are deliberately unset and the routes carry the
# explicit hostname list instead; allowedRoutes Same keeps other namespaces off
# these listeners.
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: haproxy
namespace: haproxy
labels:
traefik.io/instance: external
spec:
gatewayClassName: traefik-external
listeners:
- name: http
port: 80
protocol: HTTP
allowedRoutes:
namespaces:
from: Same
- name: https-passthrough
port: 443
protocol: TLS
tls:
mode: Passthrough
allowedRoutes:
namespaces:
from: Same
+37
View File
@@ -0,0 +1,37 @@
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: haproxy-http-redirect
namespace: haproxy
labels:
app: haproxy
spec:
hostnames:
- sonarr.main.unkin.net
- radarr.main.unkin.net
- lidarr.main.unkin.net
- readarr.main.unkin.net
- prowlarr.main.unkin.net
- nzbget.main.unkin.net
- jellyfin.main.unkin.net
- fafflix.unkin.net
- git.unkin.net
- grafana.unkin.net
- dashboard.ceph.unkin.net
- auth.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: haproxy
sectionName: http
rules:
- filters:
- type: RequestRedirect
requestRedirect:
scheme: https
statusCode: 301
matches:
- path:
type: PathPrefix
value: /
+15
View File
@@ -0,0 +1,15 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- configmap.yaml
- deployment.yaml
- service.yaml
- gateway.yaml
- tlsroute.yaml
- httproute.yaml
- pdb.yaml
- vpa.yaml
- vmpodscrape.yaml
+5
View File
@@ -0,0 +1,5 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: haproxy
+11
View File
@@ -0,0 +1,11 @@
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: haproxy
namespace: haproxy
spec:
maxUnavailable: 1
selector:
matchLabels:
app: haproxy
+19
View File
@@ -0,0 +1,19 @@
---
apiVersion: v1
kind: Service
metadata:
name: haproxy
namespace: haproxy
spec:
type: ClusterIP
# Reached only by the external Traefik's TLS-passthrough TLSRoute, so the
# peer address here is a Traefik pod, not the client. sessionAffinity is
# deliberately absent: keyed on ClientIP it would pin whole Traefik pods,
# not clients. Backend persistence rests on the per-backend SRVNAME cookie.
selector:
app: haproxy
ports:
- name: https
port: 443
protocol: TCP
targetPort: https
+34
View File
@@ -0,0 +1,34 @@
---
apiVersion: gateway.networking.k8s.io/v1
kind: TLSRoute
metadata:
name: haproxy
namespace: haproxy
labels:
app: haproxy
spec:
hostnames:
- sonarr.main.unkin.net
- radarr.main.unkin.net
- lidarr.main.unkin.net
- readarr.main.unkin.net
- prowlarr.main.unkin.net
- nzbget.main.unkin.net
- jellyfin.main.unkin.net
- fafflix.unkin.net
- git.unkin.net
- grafana.unkin.net
- dashboard.ceph.unkin.net
- auth.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: haproxy
sectionName: https-passthrough
rules:
- backendRefs:
- group: ""
kind: Service
name: haproxy
port: 443
weight: 1
+13
View File
@@ -0,0 +1,13 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: haproxy
namespace: haproxy
spec:
selector:
matchLabels:
app: haproxy
podMetricsEndpoints:
- port: metrics
path: /metrics
+13
View File
@@ -0,0 +1,13 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: haproxy-vpa
namespace: haproxy
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: haproxy
updatePolicy:
updateMode: "Off"
+3 -6
View File
@@ -1,16 +1,13 @@
--- ---
# Log ingestion endpoint for puppet-managed VMs (and any non-k8s client). # Log ingestion endpoint for puppet-managed VMs (and any non-k8s client):
# Reuses the internal Traefik gateway + cert-manager + external-dns pattern so # fronts the VLCluster vlinsert service over TLS at a name VMs can resolve.
# VMs reach the Vector aggregator's HTTP source over TLS at a DNS name they can
# resolve. The puppet-side Vector rollout ships NDJSON to
# https://logs-ingest.k8s.syd1.au.unkin.net/ (a later task).
apiVersion: gateway.networking.k8s.io/v1 apiVersion: gateway.networking.k8s.io/v1
kind: Gateway kind: Gateway
metadata: metadata:
name: logs-ingest name: logs-ingest
namespace: logging namespace: logging
labels: labels:
app.kubernetes.io/name: vector-aggregator app.kubernetes.io/name: victorialogs
app.kubernetes.io/component: ingest app.kubernetes.io/component: ingest
traefik.io/instance: internal traefik.io/instance: internal
annotations: annotations:
+4 -4
View File
@@ -5,7 +5,7 @@ metadata:
name: logs-ingest-http-redirect name: logs-ingest-http-redirect
namespace: logging namespace: logging
labels: labels:
app.kubernetes.io/name: vector-aggregator app.kubernetes.io/name: victorialogs
app.kubernetes.io/component: ingest app.kubernetes.io/component: ingest
spec: spec:
hostnames: hostnames:
@@ -32,7 +32,7 @@ metadata:
name: logs-ingest name: logs-ingest
namespace: logging namespace: logging
labels: labels:
app.kubernetes.io/name: vector-aggregator app.kubernetes.io/name: victorialogs
app.kubernetes.io/component: ingest app.kubernetes.io/component: ingest
spec: spec:
hostnames: hostnames:
@@ -46,8 +46,8 @@ spec:
- backendRefs: - backendRefs:
- group: "" - group: ""
kind: Service kind: Service
name: vector-vm-ingest name: vlinsert-logs
port: 8080 port: 9481
weight: 1 weight: 1
matches: matches:
- path: - path:
+1
View File
@@ -10,6 +10,7 @@ resources:
- job_clickhouse-schema.yaml - job_clickhouse-schema.yaml
- nats-bootstrap-job.yaml - nats-bootstrap-job.yaml
- cephrgw.yaml - cephrgw.yaml
- vlcluster.yaml
- gateway.yaml - gateway.yaml
- httproute.yaml - httproute.yaml
- serviceaccount_logarchiver.yaml - serviceaccount_logarchiver.yaml
+47
View File
@@ -0,0 +1,47 @@
---
apiVersion: operator.victoriametrics.com/v1
kind: VLCluster
metadata:
name: logs
namespace: logging
spec:
clusterVersion: v1.52.0
vlinsert:
replicaCount: 2
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
cpu: "2"
memory: 4Gi
vlselect:
replicaCount: 2
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
cpu: "2"
memory: 4Gi
vlstorage:
replicaCount: 3
retentionPeriod: 180d
# ~3 GiB/day measured; 220GiB/node cap keeps 180d time-based, not disk-bound
retentionMaxDiskSpaceUsageBytes: 220GiB
storage:
volumeClaimTemplate:
spec:
accessModes:
- ReadWriteOnce
storageClassName: cephrbd-fast-delete
resources:
requests:
storage: 250Gi
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "4"
memory: 8Gi
+1 -1
View File
@@ -25,7 +25,7 @@ spec:
- name: pdbmux - name: pdbmux
# Image is published by the pdbmux repo's .woodpecker/docker.yaml on # Image is published by the pdbmux repo's .woodpecker/docker.yaml on
# a v* tag. It only exists after that tag is cut (see PR merge gates). # a v* tag. It only exists after that tag is cut (see PR merge gates).
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/pdbmux:v0.2.0 image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/pdbmux:v0.4.0
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
ports: ports:
- containerPort: 8080 - containerPort: 8080
+2
View File
@@ -11,11 +11,13 @@ metadata:
namespace: puppet namespace: puppet
spec: spec:
schedule: "*/1 * * * *" schedule: "*/1 * * * *"
startingDeadlineSeconds: 200
concurrencyPolicy: Forbid concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3 successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3 failedJobsHistoryLimit: 3
jobTemplate: jobTemplate:
spec: spec:
activeDeadlineSeconds: 300
template: template:
metadata: metadata:
labels: labels:
@@ -99,6 +99,23 @@ spec:
- mountPath: /docker-custom-entrypoint.d/post-startup/additional-ruby-gems.sh - mountPath: /docker-custom-entrypoint.d/post-startup/additional-ruby-gems.sh
name: additional-ruby-gems name: additional-ruby-gems
subPath: additional-ruby-gems.sh subPath: additional-ruby-gems.sh
- mountPath: /configmaps/auth.conf
name: compiler-auth-conf
subPath: auth.conf
- mountPath: /docker-custom-entrypoint.d/pre-default/10-auth-conf.sh
name: compiler-auth-conf-seed
subPath: 10-auth-conf.sh
- mountPath: /docker-custom-entrypoint.d/pre-default/20-vault-helpers.sh
name: compiler-vault-helpers-seed
subPath: 20-vault-helpers.sh
- mountPath: /opt/certmanager/config.yaml
name: certmanager-config
subPath: certmanager.yaml
readOnly: true
- mountPath: /opt/sshsignhost/config.yaml
name: sshsignhost-config
subPath: sshsignhost.yaml
readOnly: true
initContainers: initContainers:
- name: copy-configmaps - name: copy-configmaps
image: busybox:1.35 image: busybox:1.35
@@ -196,7 +213,38 @@ spec:
echo "$EXPECTED encapic" | sha256sum -c - echo "$EXPECTED encapic" | sha256sum -c -
install -m 0755 encapic /opt/bin/encapic install -m 0755 encapic /opt/bin/encapic
# Puppet shells out to these two from generate() during catalog
# compilation: profiles::pki::vault runs certmanager and
# profiles::ssh::sign runs sshsignhost.
install_release() {
name=$1
version=$2
asset="$name-linux-amd64"
base="https://git.unkin.net/unkin/$name/releases/download/$version"
curl -fsSL -o "$name" "$base/$asset"
curl -fsSL -o "$name.checksums" "$base/checksums.txt"
# checksums.txt covers every release asset; pick the line for the
# one we downloaded and verify it under our local filename.
expected=$(awk -v a="$asset" '$NF == a || $NF == "*"a {print $1}' "$name.checksums")
if [ -z "$expected" ]; then
echo "no checksum for $asset in $version checksums.txt" >&2
exit 1
fi
echo "$expected $name" | sha256sum -c -
install -m 0755 "$name" "/opt/bin/$name"
}
install_release certmanager v0.2.0
install_release sshsignhost v0.1.0
echo "Shared binaries setup completed" echo "Shared binaries setup completed"
resources:
limits:
cpu: 300m
memory: 256Mi
requests:
cpu: 100m
memory: 64Mi
volumeMounts: volumeMounts:
- mountPath: /opt/bin/ - mountPath: /opt/bin/
name: puppet-shared-bins name: puppet-shared-bins
@@ -234,5 +282,22 @@ spec:
configMap: configMap:
name: additional-ruby-gems name: additional-ruby-gems
defaultMode: 0755 defaultMode: 0755
- name: compiler-auth-conf
configMap:
name: compiler-auth.conf
- name: compiler-auth-conf-seed
configMap:
name: compiler-auth-conf-seed
defaultMode: 0755
- name: compiler-vault-helpers-seed
configMap:
name: compiler-vault-helpers-seed
defaultMode: 0755
- name: certmanager-config
configMap:
name: certmanager-config
- name: sshsignhost-config
configMap:
name: sshsignhost-config
strategy: strategy:
type: RollingUpdate type: RollingUpdate
+25
View File
@@ -54,6 +54,31 @@ configMapGenerator:
- resources/compiler/puppetdb.conf - resources/compiler/puppetdb.conf
options: options:
disableNameSuffixHash: true disableNameSuffixHash: true
- name: compiler-auth.conf
files:
- resources/compiler/auth.conf
options:
disableNameSuffixHash: true
- name: compiler-auth-conf-seed
files:
- resources/compiler/10-auth-conf.sh
options:
disableNameSuffixHash: true
- name: compiler-vault-helpers-seed
files:
- resources/compiler/20-vault-helpers.sh
options:
disableNameSuffixHash: true
- name: certmanager-config
files:
- resources/compiler/certmanager.yaml
options:
disableNameSuffixHash: true
- name: sshsignhost-config
files:
- resources/compiler/sshsignhost.yaml
options:
disableNameSuffixHash: true
- name: additional-ruby-gems - name: additional-ruby-gems
files: files:
- resources/additional-ruby-gems.sh - resources/additional-ruby-gems.sh
@@ -6,4 +6,6 @@ echo "Installing additional Ruby gems..."
/opt/puppetlabs/puppet/bin/gem install ipaddr /opt/puppetlabs/puppet/bin/gem install ipaddr
/opt/puppetlabs/puppet/bin/gem install hiera-eyaml /opt/puppetlabs/puppet/bin/gem install hiera-eyaml
/opt/puppetlabs/puppet/bin/gem install toml /opt/puppetlabs/puppet/bin/gem install toml
# Under set -e a failed install kills the entrypoint post-startup hooks, taking down an already-serving compiler.
/opt/puppetlabs/bin/puppetserver gem install toml
echo "Additional Ruby gems installed successfully" echo "Additional Ruby gems installed successfully"
+14
View File
@@ -0,0 +1,14 @@
#!/bin/bash
set -euo pipefail
SRC=/configmaps/auth.conf
DST=/etc/puppetlabs/puppetserver/conf.d/auth.conf
# Copied rather than mounted: the entrypoint chowns conf.d and rewrites auth.conf,
# both of which fail on a read-only configmap mount and abort container startup.
if [ ! -s "$SRC" ]; then
echo "FATAL: $SRC missing or empty; refusing to start on the image default auth.conf" >&2
exit 1
fi
cp "$SRC" "$DST"
+29
View File
@@ -0,0 +1,29 @@
#!/bin/bash
set -euo pipefail
BIN_DIR=/opt/bin
CA=/opt/vault-ca-cert.crt
if [ ! -s "$CA" ]; then
echo "FATAL: $CA missing or empty; certmanager and sshsignhost cannot verify Vault" >&2
exit 1
fi
# profiles::pki::vault and profiles::ssh::sign shell out to fixed /usr/local/bin
# paths from generate(); the binaries ship on the shared PVC, and /usr/local/bin
# lives in the image. Wrappers rather than symlinks because neither binary reads
# a CA path from its config: SSL_CERT_FILE scopes the internal CA to these two
# processes instead of the puppetserver JVM's own trust store.
for bin in certmanager sshsignhost; do
if [ ! -x "$BIN_DIR/$bin" ]; then
echo "FATAL: $BIN_DIR/$bin missing; generate() would abort every catalog compile" >&2
exit 1
fi
cat > "/usr/local/bin/$bin" <<WRAPPER
#!/bin/sh
SSL_CERT_FILE=$CA
export SSL_CERT_FILE
exec $BIN_DIR/$bin "\$@"
WRAPPER
chmod 0755 "/usr/local/bin/$bin"
done
@@ -0,0 +1,320 @@
# Copied into conf.d at startup by 10-auth-conf.sh; the entrypoint then appends the
# admin API cache rule and re-renders the result, so the running file is not byte-identical.
authorization: {
version: 1
rules: [
{
# Allow nodes to retrieve their own catalog
match-request: {
path: "^/puppet/v3/catalog/([^/]+)$"
type: regex
method: [get, post]
}
allow: "$1"
sort-order: 500
name: "puppetlabs v3 catalog from agents"
},
{
# Allow catalog-diff to retrieve catalogs on behalf of others.
# sort-order 400 must stay lower than the puppetlabs deny that follows: rules
# sort by [sort-order, name] and the first match wins.
match-request: {
path: "^/puppet/v4/catalog/?$"
type: regex
method: post
}
allow: "catalog-diff.main.unkin.net"
sort-order: 400
name: "unkin v4 catalog for catalog-diff"
},
{
# Allow services to retrieve catalogs on behalf of others
match-request: {
path: "^/puppet/v4/catalog/?$"
type: regex
method: post
}
deny: "*"
sort-order: 500
name: "puppetlabs v4 catalog for services"
},
{
# Allow nodes to retrieve the certificate they requested earlier
match-request: {
path: "/puppet-ca/v1/certificate/"
type: path
method: get
}
allow-unauthenticated: true
sort-order: 500
name: "puppetlabs certificate"
},
{
# Allow all nodes to access the certificate revocation list
match-request: {
path: "/puppet-ca/v1/certificate_revocation_list/ca"
type: path
method: get
}
allow-unauthenticated: true
sort-order: 500
name: "puppetlabs crl"
},
{
# Allow nodes to request a new certificate
match-request: {
path: "/puppet-ca/v1/certificate_request"
type: path
method: [get, put]
}
allow-unauthenticated: true
sort-order: 500
name: "puppetlabs csr"
},
{
# Allow nodes to renew their certificate
match-request: {
path: "/puppet-ca/v1/certificate_renewal"
type: path
method: post
}
# this endpoint should never be unauthenticated, as it requires the cert to be provided.
allow: "*"
sort-order: 500
name: "puppetlabs certificate renewal"
},
{
# Allow the CA CLI to access the certificate_status endpoint
match-request: {
path: "/puppet-ca/v1/certificate_status"
type: path
method: [get, put, delete]
}
allow: {
extensions: {
pp_cli_auth: "true"
}
}
sort-order: 500
name: "puppetlabs cert status"
},
{
match-request: {
path: "^/puppet-ca/v1/certificate_revocation_list$"
type: regex
method: put
}
allow: {
extensions: {
pp_cli_auth: "true"
}
}
sort-order: 500
name: "puppetlabs CRL update"
},
{
# Allow the CA CLI to access the certificate_statuses endpoint
match-request: {
path: "/puppet-ca/v1/certificate_statuses"
type: path
method: get
}
allow: {
extensions: {
pp_cli_auth: "true"
}
}
sort-order: 500
name: "puppetlabs cert statuses"
},
{
# Allow authenticated access to the CA expirations endpoint
match-request: {
path: "/puppet-ca/v1/expirations"
type: path
method: get
}
allow: "*"
sort-order: 500
name: "puppetlabs CA cert and CRL expirations"
},
{
# Allow the CA CLI to access the certificate clean endpoint
match-request: {
path: "/puppet-ca/v1/clean"
type: path
method: put
}
allow: {
extensions: {
pp_cli_auth: "true"
}
}
sort-order: 500
name: "puppetlabs cert clean"
},
{
# Allow the CA CLI to access the certificate sign endpoint
match-request: {
path: "/puppet-ca/v1/sign"
type: path
method: post
}
allow: {
extensions: {
pp_cli_auth: "true"
}
}
sort-order: 500
name: "puppetlabs cert sign"
},
{
# Allow the CA CLI to access the certificate sign all endpoint
match-request: {
path: "/puppet-ca/v1/sign/all"
type: path
method: post
}
allow: {
extensions: {
pp_cli_auth: "true"
}
}
sort-order: 500
name: "puppetlabs cert sign all"
},
{
# Allow unauthenticated access to the status service endpoint
match-request: {
path: "/status/v1/services"
type: path
method: get
}
allow-unauthenticated: true
sort-order: 500
name: "puppetlabs status service - full"
},
{
match-request: {
path: "/status/v1/simple"
type: path
method: get
}
allow-unauthenticated: true
sort-order: 500
name: "puppetlabs status service - simple"
},
{
match-request: {
path: "/puppet/v3/environments"
type: path
method: get
}
allow: "*"
sort-order: 500
name: "puppetlabs environments"
},
{
# Allow nodes to access all file_bucket_files. Note that access for
# the 'delete' method is forbidden by Puppet regardless of the
# configuration of this rule.
match-request: {
path: "/puppet/v3/file_bucket_file"
type: path
method: [get, head, post, put]
}
allow: "*"
sort-order: 500
name: "puppetlabs file bucket file"
},
{
# Allow nodes to access all file_content. Note that access for the
# 'delete' method is forbidden by Puppet regardless of the
# configuration of this rule.
match-request: {
path: "/puppet/v3/file_content"
type: path
method: [get, post]
}
allow: "*"
sort-order: 500
name: "puppetlabs file content"
},
{
# Allow nodes to access all file_metadata. Note that access for the
# 'delete' method is forbidden by Puppet regardless of the
# configuration of this rule.
match-request: {
path: "/puppet/v3/file_metadata"
type: path
method: [get, post]
}
allow: "*"
sort-order: 500
name: "puppetlabs file metadata"
},
{
# Allow nodes to retrieve only their own node definition
match-request: {
path: "^/puppet/v3/node/([^/]+)$"
type: regex
method: get
}
allow: "$1"
sort-order: 500
name: "puppetlabs node"
},
{
# Allow nodes to store only their own reports
match-request: {
path: "^/puppet/v3/report/([^/]+)$"
type: regex
method: put
}
allow: "$1"
sort-order: 500
name: "puppetlabs report"
},
{
# Allow nodes to update their own facts
match-request: {
path: "^/puppet/v3/facts/([^/]+)$"
type: regex
method: put
}
allow: "$1"
sort-order: 500
name: "puppetlabs facts"
},
{
match-request: {
path: "/puppet/v3/static_file_content"
type: path
method: get
}
allow: "*"
sort-order: 500
name: "puppetlabs static file content"
},
{
match-request: {
path: "/puppet/v3/tasks"
type: path
}
allow: "*"
sort-order: 500
name: "puppet tasks information"
},
{
# Deny everything else. This ACL is not strictly
# necessary, but illustrates the default policy
match-request: {
path: "/"
type: path
}
deny: "*"
sort-order: 999
name: "puppetlabs deny all"
}
]
}
@@ -0,0 +1,12 @@
---
vault:
addr: https://vault.service.consul:8200
auth_method: kubernetes
k8s_mount: k8s/au/syd1
k8s_role: puppet_certmanager
jwt_path: /var/run/secrets/kubernetes.io/serviceaccount/token
mount_point: pki_int
role_name: servers_default
output_path: /tmp/certmanager
tls_skip_verify: false
timeout: 30s
@@ -0,0 +1,11 @@
---
vault:
addr: https://vault.service.consul:8200
auth_method: kubernetes
k8s_mount: k8s/au/syd1
k8s_role: puppet_sshsigner
jwt_path: /var/run/secrets/kubernetes.io/serviceaccount/token
mount_point: sshca
role_name: signhost
tls_skip_verify: false
timeout: 30s
@@ -0,0 +1,39 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: gocache-nginx
namespace: woodpecker
labels:
app.kubernetes.io/name: gocache
app.kubernetes.io/component: proxy
data:
nginx.conf: |
worker_processes auto;
error_log /dev/stderr warn;
pid /tmp/nginx.pid;
events {
worker_connections 512;
}
# GOCACHEPROG is a raw byte stream, not HTTP, so this must be stream{} not http{}.
stream {
server {
listen 9090;
# The protocol has no authentication: anyone who can reach this port can
# write cache entries, which become code in every build that reads them.
# Loopback is the kubectl port-forward fallback; in a pod netns it is
# only these two containers.
allow 127.0.0.1/32;
allow 10.10.12.200/32;
allow 10.42.0.0/16;
deny all;
# A connect session lasts the whole build; the 10m default cuts long builds off.
proxy_timeout 2h;
proxy_connect_timeout 5s;
proxy_pass 127.0.0.1:9080;
}
}
@@ -0,0 +1,131 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: gocache
namespace: woodpecker
annotations:
configmap.reloader.stakater.com/reload: "gocache-nginx"
secret.reloader.stakater.com/reload: "gocache-s3,vault-ca-cert"
labels:
app.kubernetes.io/name: gocache
app.kubernetes.io/component: cache
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/name: gocache
template:
metadata:
labels:
app.kubernetes.io/name: gocache
app.kubernetes.io/component: cache
spec:
serviceAccountName: default
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
fsGroup: 65532
seccompProfile:
type: RuntimeDefault
containers:
- name: go-cache-plugin
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/go-cache-plugin:v0.1.0
imagePullPolicy: IfNotPresent
# Root flags must precede the subcommand; only --plugin belongs to serve.
args:
- --cache-dir=/var/cache/gocache
- --bucket=gocache
# Explicit region skips the GetBucketLocation probe, which RGW handles poorly.
- --region=us-east-1
- --s3-endpoint-url=https://s3.ceph.unkin.net
- --s3-path-style
- serve
- --plugin=9080
env:
- name: AWS_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: gocache-s3
key: AWS_ACCESS_KEY_ID
- name: AWS_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: gocache-s3
key: AWS_SECRET_ACCESS_KEY
# s3.ceph.unkin.net is served by the estate CA, not a public root.
- name: AWS_CA_BUNDLE
value: /etc/ssl/vault-ca/ca.crt
volumeMounts:
- name: cache
mountPath: /var/cache/gocache
- name: vault-ca
mountPath: /etc/ssl/vault-ca
readOnly: true
securityContext:
runAsUser: 65532
runAsGroup: 65532
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: "2"
memory: 2Gi
- name: nginx
image: docker.io/nginx:1.29.8-alpine
imagePullPolicy: IfNotPresent
# Bypass the image entrypoint: its config scripts write to a read-only rootfs.
command:
- nginx
- -g
- daemon off;
ports:
- containerPort: 9090
name: gocache
protocol: TCP
volumeMounts:
- name: nginx-config
mountPath: /etc/nginx/nginx.conf
subPath: nginx.conf
readOnly: true
- name: tmp
mountPath: /tmp
securityContext:
runAsUser: 101
runAsGroup: 101
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 500m
memory: 128Mi
volumes:
# Staging cache in front of S3: losing it costs a repopulate, not data.
- name: cache
emptyDir:
sizeLimit: 20Gi
- name: nginx-config
configMap:
name: gocache-nginx
- name: tmp
emptyDir: {}
- name: vault-ca
secret:
secretName: vault-ca-cert
items:
- key: ca.crt
path: ca.crt
+33
View File
@@ -0,0 +1,33 @@
---
# Shared Go build cache (GOCACHEPROG) for CI and developer laptops. Lives in the
# woodpecker namespace because CI is the primary consumer and reads the Secret here.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: gocache
namespace: woodpecker
spec:
displayName: "Go build cache owner"
uid: gocache
maxBuckets: 1
secretName: gocache-s3
retainOnDelete: false
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: gocache
namespace: woodpecker
spec:
bucketName: gocache
ownerRef: gocache
versioning: false
# No placementTarget: default (replicated) placement, not the ec target the
# backup buckets use — a build cache is millions of small objects.
tags:
app: gocache
purpose: go-build-cache
retainOnDelete: false
# A cache bucket is never empty, and the operator refuses to delete a
# non-empty bucket without this, wedging the finalizer.
purgeOnDelete: true
+5
View File
@@ -7,6 +7,10 @@ resources:
- cnpg_cluster.yaml - cnpg_cluster.yaml
- cnpg_backup.yaml - cnpg_backup.yaml
- cnpg_pooler.yaml - cnpg_pooler.yaml
- gocache_bucket.yaml
- configmap_gocache-nginx.yaml
- deployment_gocache.yaml
- service_gocache.yaml
- serviceaccount_arrproxy_ci.yaml - serviceaccount_arrproxy_ci.yaml
- serviceaccount_autobackup_operator_ci.yaml - serviceaccount_autobackup_operator_ci.yaml
- serviceaccount_ghp.yaml - serviceaccount_ghp.yaml
@@ -15,6 +19,7 @@ resources:
- serviceaccount_mediamark_ci.yaml - serviceaccount_mediamark_ci.yaml
- serviceaccount_plugin_docker_buildx.yaml - serviceaccount_plugin_docker_buildx.yaml
- serviceaccount_jellyfin_ha_src.yaml - serviceaccount_jellyfin_ha_src.yaml
- serviceaccount_jellyfin_plugin_sso.yaml
- serviceaccount_repospawner_ci.yaml - serviceaccount_repospawner_ci.yaml
- serviceaccount_terraform_artifactapi.yaml - serviceaccount_terraform_artifactapi.yaml
- serviceaccount_terraform_authentik.yaml - serviceaccount_terraform_authentik.yaml
+23
View File
@@ -0,0 +1,23 @@
---
apiVersion: v1
kind: Service
metadata:
name: gocache
namespace: woodpecker
annotations:
purelb.io/addresses: 198.18.200.11
purelb.io/service-group: common
labels:
app.kubernetes.io/name: gocache
spec:
type: LoadBalancer
# Cluster SNATs off-node traffic to a node address, which would defeat the
# nginx allow rules; Local preserves the wireguard client IP.
externalTrafficPolicy: Local
selector:
app.kubernetes.io/name: gocache
ports:
- name: gocache
port: 9090
targetPort: gocache
protocol: TCP
@@ -0,0 +1,6 @@
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: jellyfin-plugin-sso
namespace: woodpecker
@@ -0,0 +1,6 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ../../../base/haproxy
@@ -10,7 +10,7 @@ resources:
helmCharts: helmCharts:
- name: victoria-metrics-operator - name: victoria-metrics-operator
repo: https://victoriametrics.github.io/helm-charts/ repo: https://victoriametrics.github.io/helm-charts/
version: "0.57.1" version: "0.67.3"
releaseName: victoria-metrics-operator releaseName: victoria-metrics-operator
namespace: vm-system namespace: vm-system
valuesFile: values.yaml valuesFile: values.yaml
+1 -1
View File
@@ -4,7 +4,7 @@ agent:
WOODPECKER_MAX_WORKFLOWS: "8" WOODPECKER_MAX_WORKFLOWS: "8"
WOODPECKER_BACKEND_K8S_PRIORITY_CLASS: power WOODPECKER_BACKEND_K8S_PRIORITY_CLASS: power
WOODPECKER_BACKEND_K8S_STORAGE_CLASS: cephrbd-fast-delete WOODPECKER_BACKEND_K8S_STORAGE_CLASS: cephrbd-fast-delete
WOODPECKER_BACKEND_K8S_VOLUME_SIZE: 10G WOODPECKER_BACKEND_K8S_VOLUME_SIZE: 20Gi
WOODPECKER_BACKEND_K8S_STORAGE_RWX: false WOODPECKER_BACKEND_K8S_STORAGE_RWX: false
# Required from woodpecker 3.16.0 (GHSA-qf34-295c-26v8): step-level # Required from woodpecker 3.16.0 (GHSA-qf34-295c-26v8): step-level
# serviceAccountName is gated behind this agent flag (default false). # serviceAccountName is gated behind this agent flag (default false).
+1
View File
@@ -29,6 +29,7 @@ spec:
- path: apps/overlays/*/ghp - path: apps/overlays/*/ghp
- path: apps/overlays/*/gitea - path: apps/overlays/*/gitea
- path: apps/overlays/*/grafana-system - path: apps/overlays/*/grafana-system
- path: apps/overlays/*/haproxy
- path: apps/overlays/*/inteldeviceplugins-system - path: apps/overlays/*/inteldeviceplugins-system
- path: apps/overlays/*/jfrog - path: apps/overlays/*/jfrog
- path: apps/overlays/*/k8up-system - path: apps/overlays/*/k8up-system
+2
View File
@@ -43,6 +43,8 @@ spec:
server: https://kubernetes.default.svc server: https://kubernetes.default.svc
- namespace: 'gitea' - namespace: 'gitea'
server: https://kubernetes.default.svc server: https://kubernetes.default.svc
- namespace: 'haproxy'
server: https://kubernetes.default.svc
- namespace: 'jfrog' - namespace: 'jfrog'
server: https://kubernetes.default.svc server: https://kubernetes.default.svc
- namespace: 'kanidm' - namespace: 'kanidm'
@@ -26,6 +26,10 @@ data:
issuer: https://identity.unkin.net/application/o/argocd/ issuer: https://identity.unkin.net/application/o/argocd/
clientID: argocd clientID: argocd
clientSecret: $argocd-oidc:client_secret clientSecret: $argocd-oidc:client_secret
# The Authentik client is public (the iOS app can't hold a secret), so
# Authentik no longer enforces clientSecret; PKCE replaces it as the
# protection against authorization-code interception.
enablePKCEAuthentication: true
# identity.unkin.net now serves the LetsEncrypt *.unkin.net wildcard, so the # identity.unkin.net now serves the LetsEncrypt *.unkin.net wildcard, so the
# stock image trust store validates it; no rootCA pin. # stock image trust store validates it; no rootCA pin.
requestedScopes: requestedScopes: