Commit Graph

511 Commits

Author SHA1 Message Date
unkin-agent 633df4d441 bump artifactapi to v3.13.1 from docker-internal
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
2026-10-09 23:21:09 +11:00
unkin-agent 7960ee03f9 pin artifactapi back to v3.12.0 (#532)
The v3.13.0 images were never published (registry push failed), so the current pin cannot be pulled.

- pin api and ui images back to v3.12.0

Reviewed-on: #532
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-09 22:52:29 +11:00
unkin-agent 393100430b Bump artifactapi to v3.13.0 (#529)
Bump artifactapi API and UI images from v3.12.0 to v3.13.0.

- Update git.unkin.net/unkin/artifactapi to v3.13.0
- Update git.unkin.net/unkin/artifactapi-ui to v3.13.0

Reviewed-on: #529
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-09 22:48:05 +11:00
unkin-agent 5024945c74 consul: remove standalone k8s consul (#530)
The standalone k8s Consul runs its own raft under DC `au-syd1`, the same DC name as the VM cluster that k8s servers are about to join. Its leftover state risks a split brain, so it goes before the replacement lands.

- remove `apps/overlays/au-syd1/consul`, dropping `platform-consul` from the platform ApplicationSet and cascading deletion of the `consul` namespace and PVCs
- keep `apps/base/consul` for reuse by the replacement

Reviewed-on: #530
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-09 22:46:40 +11:00
unkin-agent 87880ac8ef kea: bump to v0.1.7 from docker-internal (#527)
v0.1.7 stops the operator reconcile hot-loop, which was hot-reloading Kea about 9 times a second and flooding the dhcp-system logs. Images now publish to docker-internal.

- bump kea-operator, kea and kea-api to v0.1.7
- pull all three from artifactapi docker-internal

Reviewed-on: #527
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-09 21:49:33 +11:00
unkin-agent 7ac84461ce consul: pull chart via artifactapi helm mirror (#528)
The consul overlay fetches its chart directly from helm.releases.hashicorp.com, while the other 15 helm overlays on the estate mirror use the artifactapi virtual helm repo, which already proxies hashicorp-helm.

- point consul helmCharts repo at the artifactapi virtual helm repo

Reviewed-on: #528
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-09 21:49:13 +11:00
unkin-agent 9b857bef64 logging: expose vlselect on internal traefik (#525)
Querying the VLCluster from outside the cluster (CLI tools, puppet VMs) needs a resolvable TLS endpoint for vlselect, mirroring the existing logs-ingest endpoint for vlinsert.

- add `vlselect` Gateway on `traefik-internal` with vault-issuer TLS and external-dns for `vlselect.k8s.syd1.au.unkin.net`
- add HTTP->HTTPS redirect route and HTTPS route to `vlselect-logs:9471`, unauthenticated

Reviewed-on: #525
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-05 14:08:33 +11:00
unkin-agent d1825bb655 authentik: raise server cpu request to 500m (#526)
authentik-server requests only 50m CPU, so on busy nodes it is starved, `/-/health/live/` exceeds the 3s liveness timeout and kubelet restarts it (~100 restarts per pod).

- raise `server.resources.requests.cpu` from 50m to 500m

Reviewed-on: #526
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-05 14:07:45 +11:00
unkin-agent 206fe90f80 Raise puppetserver compiler requests to match real usage (#524)
Compiler pods average ~2Gi memory but request 1Gi, so the scheduler overpacks nodes and the HPA's utilisation target is skewed.

- raise compiler requests to 1 cpu / 2Gi and limits to 4 cpu / 4Gi
- scale up faster (30s window, 75%) and down slower (600s window)

Reviewed-on: #524
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-05 12:34:19 +11:00
unkin-agent 6d842c7d66 feat(bind-internal): publish haproxy edge sites from k8s DNS (#523)
The VM haproxies are being decommissioned; the sites they published via puppet CNAMEs are already served by the k8s haproxy edge, so k8s DNS publishes them instead.

- add main.unkin.net A records for sonarr, radarr, lidarr, readarr, prowlarr, nzbget, jellyfin -> 198.18.199.0
- add unkin.net A records for fafflix and git (git-edge.yaml) -> 198.18.199.0, auth -> 198.18.200.4
- point the commented git.yaml cutover at replacing git-edge.yaml

Merge alongside puppet-prod halb DNS removal; A records only take effect once the puppet CNAMEs are gone.

Reviewed-on: #523
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-05 01:31:06 +11:00
unkin-agent d9cc24dbdb Serve grafana.unkin.net from traefik-external (#522)
grafana.unkin.net still routes through the puppet haproxy edge to the old grafana VMs; the k8s grafana should serve it directly like identity and vlogs.

- add grafana-external Gateway (traefik-external, *.unkin.net wildcard) with redirect + main HTTPRoutes
- reflect wildcard-unkin-net-tls into grafana
- set grafana root_url to https://grafana.unkin.net
- add grafana A record -> 198.18.199.0 in the bind-operator unkin.net zone
- drop grafana.unkin.net from the k8s haproxy routes and config

Requires terraform-authentik grafana redirect URI PR applied first, and the puppet halb vrrp_cnames grafana.unkin.net CNAME removed.

Reviewed-on: #522
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-05 00:38:21 +11:00
unkin-agent 115cdd492f Bump artifactapi to v3.12.0 (#521)
artifactapi v3.12.0 makes DELETE eviction clear every cache layer, so stale mutable indexes (e.g. epel repomd) no longer survive eviction. It also adds `?prefix=` filtering to object listings.

- bump artifactapi and artifactapi-ui images to v3.12.0

Reviewed-on: #521
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-04 23:41:15 +11:00
unkin-agent 9375052d6c Set traefik resource requests so HPAs can scale (#520)
Both traefik HPAs report `ScalingActive False: missing request for cpu` because the chart values set no container resources, so the CPU-utilisation metric can't be computed and neither Deployment ever scales.

- traefik-external requests 200m CPU / 128Mi, memory limit 512Mi
- traefik-internal requests 500m CPU / 256Mi, memory limit 1Gi
- No CPU limits, to avoid throttling the proxy
- The chart derives GOMEMLIMIT (90% of the memory limit) from the new limits

Reviewed-on: #520
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-04 15:27:49 +11:00
unkin-agent af61ac5e92 Set traefik-external externalTrafficPolicy to Local (#519)
With the default Cluster external policy, kube-proxy SNATs inbound traffic to the traefik-external LoadBalancer, hiding real client IPs from traefik and adding a cross-node hop. Local preserves source IPs. Internal policy stays Cluster so in-cluster callers on nodes without a traefik-external pod still reach it.

- Set `externalTrafficPolicy: Local` on the traefik-external Service
- Set `internalTrafficPolicy: Cluster` explicitly

Reviewed-on: #519
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-04 15:26:23 +11:00
unkin-agent 2bcce4894f Drop git.unkin.net from k8s gitea gateway and routes (#518)
git.unkin.net is served by the VM haproxy during the forge migration. Claiming it on the k8s gitea Gateway/HTTPRoutes competes with that path, so the k8s gitea serves only its admin name until cutover.

- remove git.unkin.net from the gitea HTTPRoute and redirect route hostnames
- remove the http-primary/https-primary listeners and their parentRefs
- set the gateway cert common-name to git.k8s.syd1.au.unkin.net

Reviewed-on: #518
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-04 15:07:08 +11:00
unkin-agent b0bfa4605f vmagent: raise kubernetes-apiservers max_scrape_size to 64MiB (#517)
The kubernetes-apiservers job reports up=0 on every target because the apiserver /metrics response exceeds vmagent's default 16MiB scrape size limit.

- set max_scrape_size: 64MiB on the kubernetes-apiservers job only

Reviewed-on: #517
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-04 14:01:14 +11:00
unkin-agent cc0dd24991 vmagent: fix consul target address relabel (#516)
The consul job joins node and port with ":" but then sets `replacement: "${1}:${2}"` with the default `(.*)` regex, so `__address__` becomes `node:port:`. vmagent rejects every consul target as an invalid port, including node_exporter.

- drop the replacement so the default `$1` yields `node:port`

Reviewed-on: #516
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-03 22:51:46 +10:00
unkin-agent 27dd7e7f3d vmagent: cap consul SD wait time under nginx timeout (#515)
nginx in front of Consul closes requests at 90s. vmagent's consul SD sends blocking queries with a 525s wait, so every watch fails with a 504 and the consul job finds no targets. That leaves no puppet-host metrics (node_exporter etc.) in VictoriaMetrics.

- set `promscrape.consul.waitTime: 50s` on the main VMAgent

Reviewed-on: #515
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-03 22:15:16 +10:00
unkin-agent 96624a69c4 Rename the terraform-infra CI ServiceAccount to terraform-netbox (#514)
terraform-infra is replaced by terraform-netbox, whose Vault kubernetes auth role binds the terraform-netbox service account in the woodpecker namespace.

- rename serviceaccount_terraform_infra to serviceaccount_terraform_netbox

Pairs with terraform-vault #161 (the auth role) and terraform-git #105 (the repo).

Reviewed-on: #514
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-03 09:34:19 +10:00
unkin-agent 73cb3903c2 Add k8up Grafana dashboard (#513)
With k8up operator metrics now scraped into VictoriaMetrics, backup job and schedule state still has no dashboard to view it.

- add the k8up dashboard from M0NsTeRRR/grafana-dashboards as a gzipped GrafanaDashboard CR
- resolve its datasource to the in-cluster VictoriaMetrics uid
- register it in the grafana kustomization

Reviewed-on: #513
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-02 23:29:40 +10:00
unkin-agent 575bd5d361 Scrape k8up operator metrics (#512)
The k8up operator already serves Prometheus metrics (backup/check/prune job counts, controller-runtime) on its k8up-metrics Service, but nothing scrapes them, so backup health is invisible in VictoriaMetrics.

- add a VMServiceScrape for the k8up-metrics Service (port http, /metrics)

Reviewed-on: #512
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-02 23:28:59 +10:00
unkin-agent 79e8ccb3d8 feat(grafana): sync kubernetes dashboards from victoria-metrics-k8s-stack (#511)
Grafana lacks Kubernetes cluster/workload dashboards; the victoria-metrics-k8s-stack sync-job ships a maintained set as GrafanaDashboard CRs.

- render victoria-metrics-k8s-stack 0.95.0 in the grafana overlay with only the dashboard sync-job enabled
- sync dotdc kubernetes-views/api-server and kube-prometheus dashboards via the artifactapi github_user remote
- skip dashboards already in apps/base/grafana and those without scraped metrics
- trust vault-ca-cert for artifactapi TLS

Requires terraform-artifactapi #49 applied first.

Reviewed-on: #511
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-02 23:04:57 +10:00
unkin-agent 5a8559e97c grafana: drop alerting Service and POD_IP env owned by grafana-operator (#510)
grafana-operator already creates the headless `grafana-alerting` Service (port 9094) and injects `POD_IP` into the grafana container, so the copies added in #509 make Argo CD and the operator fight over the same objects.

- remove `service-alerting.yaml` from the grafana base
- remove the `POD_IP` env entry from the Grafana CR
- keep the `unified_alerting` `ha_*` config, which targets the operator's Service

Reviewed-on: #510
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-02 22:44:54 +10:00
unkin-agent ec60a8c9d7 grafana: enable unified alerting HA gossip (#509)
Grafana runs 3 replicas with no alerting HA config, so every replica sends every notification.

- add headless `grafana-alerting` Service exposing gossip on 9094 TCP/UDP
- inject `POD_IP` from `status.podIP` into the grafana container
- set `unified_alerting` `ha_peers`, `ha_listen_address` and `ha_advertise_address` so replicas form one Alertmanager cluster

Reviewed-on: #509
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-02 22:37:19 +10:00
unkin-agent ee2476fe5d grafana: scale to 3 replicas, raise limits (#508)
A single Grafana pod with a 1 cpu/1Gi limit is a single point of failure and gets throttled under dashboard load. State lives in Postgres, so extra replicas are safe; spreading them keeps a node loss from taking out Grafana.

- set `replicas: 3` on the grafana deployment
- raise grafana container limits to 2 cpu / 4Gi
- spread grafana pods across nodes with a soft hostname topology constraint

Reviewed-on: #508
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-10-02 22:15:25 +10:00
unkin-agent e6e882abfc Give vlogs its own namespace (#507)
The Vault KV layout is kubernetes/namespace/<ns>/<sa>/<secret>, so vlogs and logviewer both running as SA default in the shared logging namespace would collide on one oauth-credentials entry. Splitting vlogs out resolves it without widening any Vault policy.

- Move apps/base/logging/vlogs to apps/base/vlogs, namespace vlogs
- Add namespace.yaml and a vlogs-scoped VaultAuth (role default)
- Point the VaultStaticSecret at kubernetes/namespace/vlogs/default/oauth-credentials
- Add the au-syd1 overlay, platform ApplicationSet path and project destination
- Move the wildcard-unkin-net-tls reflection from logging to vlogs; vlogs was its only consumer there

Secret is already seeded at the new Vault path.

Reviewed-on: #507
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-28 22:55:28 +10:00
unkin-agent c979579c50 Collect pod logs with VLAgent, retire the NATS/ClickHouse chain (#506)
The vector -> JetStream -> ClickHouse chain duplicated storage the in-cluster VictoriaLogs already provides, and its NATS StatefulSet keeps the logging Application from syncing.

- collect pod logs with a VLAgent DaemonSet remote-writing to vlinsert-logs
- delete vector, NATS, ClickHouse, logarchiver and logviewer
- drop clickhouse-system from the logging ApplicationSet

---------

Co-authored-by: BenVincent <benvin@main.unkin.net>
Reviewed-on: #506
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 20:23:58 +10:00
unkin-agent 9f6019b0e4 Drop obsolete vector-test CI pipeline (#505)
Vector is being removed from the logging stack in favour of a VictoriaLogs VLAgent, so every path this pipeline tests is going away and the pipeline becomes dead weight.

- Delete `.woodpecker/vector-test.yaml`.

Companion PR removes the vector manifests under `apps/base/logging/vector/`.

Reviewed-on: #505
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 20:12:45 +10:00
unkinben a5a5cc44e9 Promote bind-operator to v0.3.1 (#503)
bind-operator v0.3.1 fixes the apex NS bug two ways: the operator no longer glues apex NS records to pod IPs, and a new CEL rule on `DNSRecordSpec` rejects apex NS DNSRecords at admission. The CRDs are pinned by tag URL, so bumping only the image would leave the admission half inert.

- Bump the bind-operator image to `v0.3.1`
- Bump the CRD install URL to the matching `v0.3.1` tag

---------

Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Reviewed-on: #503
Co-authored-by: Ben Vincent <ben@unkin.net>
Co-committed-by: Ben Vincent <ben@unkin.net>
2026-09-27 18:11:38 +10:00
unkin-agent be2c751175 Redirect vlogs root to the VictoriaLogs UI (#504)
Visiting https://vlogs.unkin.net/ currently lands on the bare VictoriaLogs index page with relative links, not the query UI.

- Add an Exact `/` rule to the `vlogs` HTTPRoute that 302s to `/select/vmui/`
- Keep the PathPrefix `/` catch-all pointed at `vlogs-oauth2:80`, so auth is unchanged

Reviewed-on: #504
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 18:10:35 +10:00
unkin-agent 216b1d72ac Split bind-internal DNSRecords into one file per record (#501)
records.yaml had grown to ten DNSRecord documents across two zones, so finding or reviewing a single record meant scanning the whole file and every change touched it.

- move each record to authoritative/<zone>/<type>/<record>.yaml
- add a kustomization.yaml per zone and per type directory
- keep the git.unkin.net cutover record commented out, alongside a commented kustomization entry
- move the DNSRecord-namespace rationale onto the authoritative kustomization
- delete records.yaml

Rendered output is unchanged.

Reviewed-on: #501
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 17:43:16 +10:00
unkin-agent 7581578df1 Split bind-external DNSRecords into one file per record (#500)
A single records.yaml holding every DNSRecord means any record change touches a shared file, and deleting one record is a hunk edit rather than a file removal. One file per record under <zone>/<type>/<record>.yaml makes each record independently editable and removable.

- move acme-apex-ns to acme-unkin-net/ns/apex.yaml and acme-ns1-a to acme-unkin-net/a/ns1.yaml
- add kustomization.yaml per zone and type directory, referencing directories from the parent
- carry the delegation rationale onto the zone kustomization
- drop records.yaml and reference the zone directory from the app kustomization

Rendered output is unchanged.

Reviewed-on: #500
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 17:28:37 +10:00
unkin-agent 6515e7f637 Expose VictoriaLogs UI at vlogs.unkin.net behind oauth2-proxy (#498)
The vlselect query UI is only reachable in-cluster, so every log search needs a port-forward. Publish it at vlogs.unkin.net behind oauth2-proxy on the external (DMZ) Traefik.

- Add `apps/base/logging/vlogs`: external Gateway, http->https redirect and main HTTPRoute
- Terminate TLS with the reflected Let's Encrypt `*.unkin.net` wildcard; reflect it into `logging`
- Publish the `vlogs` A record at the DMZ gateway VIP from the bind-operator `unkin.net` zone
- Route all traffic through the `vlogs-oauth2` Service, upstreaming to `vlselect-logs:9471`
- Gate on the Authentik `vlogs` application, group `akP-vlogs-admin`
- Read OIDC credentials from `kv/kubernetes/namespace/logging/default/vlogs-oauth-credentials`

Depends on unkin/terraform-authentik#40.

Reviewed-on: #498
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 17:21:40 +10:00
unkin-agent 90d512fadf Drop the gocache serve Deployment and nginx sidecar (#499)
`connect` mode gives the client no local cache, so every cache operation is a network round trip -- about 2.7x slower than pointing the plugin straight at S3. Developers and CI run the plugin in direct mode instead.

- Remove the `go-cache-plugin serve` Deployment and its nginx stream-proxy ConfigMap
- Remove the PureLB Service, freeing `198.18.200.11`
- Keep the `gocache` bucket and `gocache-s3` Secret for direct-mode clients

Reviewed-on: #499
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 11:20:59 +10:00
unkin-agent 757ae5b240 Add a VictoriaLogs cluster and point logs-ingest at it (#488)
The k8s log pipeline stores to ClickHouse via NATS+vector, while the VM estate ships journald to a separate puppet-managed VictoriaLogs cluster. Consolidating on VictoriaLogs in-cluster collapses the two paths, and the logs-ingest gateway has no clients yet so it can be repointed now, ahead of the puppet change.

- add VLCluster `logs` at v1.52.0 (2 vlinsert, 2 vlselect, 3 vlstorage, 180d retention, 250Gi each on cephrbd-fast-delete)
- cap vlstorage disk use at 220GiB per node so 180d stays time-based rather than disk-bound
- repoint the logs-ingest HTTPRoute at `vlinsert-logs:9481`
- add a VictoriaLogs Grafana datasource and install its plugin

Nothing is removed here; NATS, ClickHouse, vector, logarchiver and logviewer keep running until a follow-up drops them.

Reviewed-on: #488
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:33:58 +10:00
unkin-agent 02f877540c Add gocache serve Deployment with nginx stream sidecar (#496)
`go-cache-plugin serve` binds `127.0.0.1` only, so nothing outside the pod can reach it and laptops have no way to use the S3-backed Go cache without holding RGW credentials.

- Run `go-cache-plugin serve` against the `gocache` bucket, path-style, explicit region to skip the GetBucketLocation probe
- Add an nginx sidecar stream-proxying `:9090` to the loopback plugin port, `proxy_timeout 2h`
- Publish it on PureLB `198.18.200.11`, `externalTrafficPolicy: Local` so the client IP reaches the allow rules
- Restrict to workstation + pod CIDRs: GOCACHEPROG is unauthenticated and a poisoned entry runs in every consuming build

Merge only after `docker-internal/go-cache-plugin:v0.1.0` is published.

Reviewed-on: #496
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:32:32 +10:00
unkin-agent 148dac8ca2 Declare the acme.unkin.net nameservers (#495)
The zone was seeded with an apex `NS ns1.acme.unkin.net` glued to the primary pod IP. Both were later corrected by hand, so the live RRset and the ns1 address exist only in the zone journal -- a reseed republishes the pod IP and breaks DNS-01 for every `*.unkin.net` cert. Declaring them makes git the source of truth.

- declare the two published apex NS names
- declare the in-zone ns1 address, which a seed would otherwise glue to the pod IP

Matches what the zone serves today, so applying it changes no records. Requires bind-operator v0.3.0.

Reviewed-on: #495
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:10:13 +10:00
unkin-agent 53e5846c18 Roll bind-operator to v0.3.0 (#494)
v0.3.0 converges a zone apex NS onto its declared nameservers instead of leaving the seed placeholder, which publishes a primary pod IP. The CRD moves with the image because the operator reads the new `spec.nameservers` field.

- pin the bind-operator image to v0.3.0
- pull the CRDs from the v0.3.0 tag

Reviewed-on: #494
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:06:21 +10:00
unkin-agent 9fed5decc8 bump victoria-metrics-operator chart to 0.67.3 (#493)
The vm-system overlay pins victoria-metrics-operator chart 0.57.1 (operator v0.66.1), eight operator minors behind upstream, so the cluster runs without newer CRD fields and reconciler fixes.

- bump the victoria-metrics-operator helmChart version to 0.67.3 (operator v0.74.1)

Reviewed-on: #493
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:05:47 +10:00
unkin-agent 20077f1029 Move the haproxy edge behind the external Traefik (#492)
The haproxy edge holds its own DMZ VIP, a second public entry point alongside
traefik-external that must be firewalled and DNS'd separately. Traefik can
front it with TLS passthrough, leaving haproxy's certs and backends untouched.

- Add a `traefik-external` Gateway: HTTP :80 plus Passthrough TLS :443.
- TLSRoute the 12 `fe_https.map` hostnames to haproxy:443; HTTPRoute 301s :80.
- Make the Service ClusterIP on 443 only, releasing 198.18.199.1.
- Drop `fe_http`, `be_letsencrypt` and `fe_http.map`; certs are DNS-01 only.

Client IP now reads as a Traefik pod — the Gateway provider cannot emit PROXY protocol to a TLSRoute backend. `sessionAffinity` goes too (it would pin Traefik pods, not clients); SRVNAME cookies keep persistence.

Reviewed-on: #492
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 23:38:07 +10:00
unkin-agent 426a399f31 Drop stalwart mail proxying from the haproxy edge (#491)
Stalwart was only ever a test deployment. The daemon is dead on all three
backend VMs and nothing public depends on it — `unkin.net` MX points at Google —
so the edge is proxying mail to nowhere and the tcp frontends make `defaults`
emit 20 spurious HTTP-mode warnings.

- Drop the `fe_smtp`, `fe_submission`, `fe_imap` and `fe_imaps` frontends.
- Drop the five `be_stalwart_*` backends and their map entries in `fe_http.map`/`fe_https.map`.
- Drop the now-unused 25/143/587/993 Service and container ports.

`haproxy -c` on the rendered config: exit 0, 0 warnings (was 20), 0 alerts.

Reviewed-on: #491
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 21:25:49 +10:00
unkin-agent 5341253573 Add Ceph RGW bucket for the shared Go build cache (#489)
Go builds on CI and laptops each rebuild the same packages from scratch. A
GOCACHEPROG backend needs an S3 bucket plus credentials before anything can
point at it, so provision those first. The bucket lives in the woodpecker
namespace because CI is the primary consumer and reads the Secret there.

- add Bucket and ObjectStoreUser for the shared Go build cache
- use default (replicated) placement rather than the ec target, since a build
  cache is millions of small objects
- purge and drop the bucket and user on delete; the cache is disposable

Nothing consumes the bucket yet.

Reviewed-on: #489
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:56:51 +10:00
unkin-agent d48125d699 Add job and start deadlines to the g10k-code CronJob (#490)
A g10k-code job wedged in ContainerCreating on a failed CephFS mount and never reached a terminal condition, so it stayed in the CronJob active list and `concurrencyPolicy: Forbid` skipped every following minute. No Puppet code reached the estate for 6 days, and the piled-up missed slots crossed the controller 100-slot cap into `TooManyMissedTimes`. The CronJob carried no deadlines at all.

- Cap a job at `activeDeadlineSeconds: 300` on the Job spec, so a hang is failed as `DeadlineExceeded` and drops out of the active list (healthy runs take 16-18s).
- Set `startingDeadlineSeconds: 200`, bounding missed-schedule look-back to ~3 slots so the count cannot reach 100.

Reviewed-on: #490
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:55:32 +10:00
unkin-agent abf6bfae88 Drop dead X-Frame-Options rules from the haproxy edge (#487)
The 13 `X-Frame-Options DENY if acl_<host>` rules in `fe_https` have never fired:
their ACLs use `req.hdr(host)`, a request-direction fetch that is invalid in a
response ruleset, so HAProxy rejects them at config-check time. Carried over
verbatim from the Puppet LXD config during the k8s move.

- Remove the 13 dead `http-response set-header X-Frame-Options` rules.
- Remove the 15 now-orphaned `acl acl_*` definition lines.

Not switching the header on: it has never been live, and Grafana/Gitea send their
own. `haproxy -c` warnings drop 33 -> 20; the two working response headers stay.

Reviewed-on: #487
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:36:26 +10:00
unkin-agent 4190785389 Move the au-syd1 haproxy edge into Kubernetes (#485)
The au-syd1 edge proxy runs on a hand-managed LXD container outside the cluster, with no HA and no shared config source.

- Add `apps/base/haproxy/`: 3 replicas behind the DMZ LoadBalancer 198.18.199.1, config from a ConfigMap, wildcard certs from reflected secrets.
- Keep source IPs via `externalTrafficPolicy: Local`; `sessionAffinity: ClientIP` stands in for the stick-table peers a Deployment cannot name.
- Drain on shutdown: `hard-stop-after 2m`, a preStop SIGUSR1 soft-stop, 150s grace.
- Bind the stats listener to 127.0.0.1 so it is port-forward only.
- Register the app in the platform project and ApplicationSet.

Reviewed-on: #485
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 18:22:57 +10:00
unkin-agent 812a9a2f2b Publish real delegation records for acme.unkin.net (#486)
The acme.unkin.net zone still serves only the bind-operator seed apex: NS ns1.acme.unkin.net glued to A 10.42.6.38, a pod IP no pod holds. The parent delegates to acme-ns1.unkin.net, but public resolvers have already promoted the child NS RRset, so when the cached address expires DNS-01 fails for every unkin.net wildcard at once.

- Add apex NS acme-ns1.unkin.net., matching the parent delegation (out of zone, no glue needed).
- Point ns1.acme.unkin.net at 103.216.191.185 so resolvers holding the seeded NS name still reach the zone.
- The operator seed placeholder itself is tracked separately in bind-operator.

---------

Co-authored-by: unkin-agent <agent@unkin.net>
Reviewed-on: #486
Co-authored-by: Unkin Agent <unkin-agent@unkin.net>
Co-committed-by: Unkin Agent <unkin-agent@unkin.net>
2026-09-26 16:42:02 +10:00
unkin-agent f37749523d Add *.main and *.ceph wildcard certificates for haproxy (#484)
The haproxy edge terminates TLS for hosts under `main.unkin.net` and `ceph.unkin.net`, which the single `*.unkin.net` wildcard does not cover.

- Add cert-manager Certificates for both wildcards from the `letsencrypt` ClusterIssuer.
- Reflect the minted secrets into the `haproxy` namespace.

Needs these records in the public unkin.net zone first:
`_acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.`
`_acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.`

Reviewed-on: #484
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 16:38:31 +10:00
unkin-agent fe51aa07be Give the puppetserver compilers the Vault cert helpers (#482)
profiles::pki::vault and profiles::ssh::sign shell out to
/usr/local/bin/certmanager and /usr/local/bin/sshsignhost from generate()
during catalog compilation. Neither binary exists in the compiler image, so
every node using them fails to compile.

- install certmanager v0.2.0 and sshsignhost v0.1.0 onto the shared bin volume with sha256 verification
- wrap both at /usr/local/bin from a pre-default entrypoint hook, failing startup loudly if either is missing
- mount read-only Vault configs for both: kubernetes auth on k8s/au/syd1, internal CA verified rather than skipped

Reviewed-on: #482
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-24 21:15:09 +10:00
unkin-agent cdaab736b5 Bump jellyfin-ha to v0.4.0 (#483)
The deployed v0.3.3 build returns 500 from /Shows/NextUp on PostgreSQL, breaking the home screen, and lets replicas diverge: library-visibility and shared-config changes never propagate, user data (resume, played state, favourites, ratings) is overwritten between pods, and eight scheduled tasks run on every replica instead of only the scan leader. v0.4.0 carries the fixes.

- Pin cheeztv and fafflix to jellyfin-ha:v0.4.0.

No config change needed: cross-pod invalidation reuses the transcode-store Redis connection string both apps already set.

Reviewed-on: #483
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-22 22:57:06 +10:00
unkin-agent b31517e6d9 Merge pull request #481 from benvin/jellyfin-sso-valkey-state
Roll jellyfin-ha to v0.3.3 and drop Service session affinity

Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 14:02:48 +10:00