A ClusterIP is only reachable via kubectl port-forward, which needs a
create grant on pods/portforward that the read-only operator context does
not have; laptops reach the cache over wireguard.
- Publish the Service as a LoadBalancer on 198.18.200.11 in the common pool
- Set externalTrafficPolicy Local so the client IP survives to the nginx
allow rules, matching the bind LoadBalancers
go-cache-plugin serve binds 127.0.0.1 only, so a Service cannot reach it
directly and CI/developer builds have no way to use the S3-backed cache.
- Run go-cache-plugin serve against the gocache RGW bucket, path-style, with
an explicit region to skip the GetBucketLocation probe
- Add an nginx sidecar stream-proxying 9090 to the loopback plugin port
- Restrict the listener to the workstation and pod CIDRs: GOCACHEPROG is
unauthenticated and a poisoned entry runs in every consuming build
- Stage the cache on an emptyDir; loss costs a repopulate from S3
The zone was seeded with an apex `NS ns1.acme.unkin.net` glued to the primary pod IP. Both were later corrected by hand, so the live RRset and the ns1 address exist only in the zone journal -- a reseed republishes the pod IP and breaks DNS-01 for every `*.unkin.net` cert. Declaring them makes git the source of truth.
- declare the two published apex NS names
- declare the in-zone ns1 address, which a seed would otherwise glue to the pod IP
Matches what the zone serves today, so applying it changes no records. Requires bind-operator v0.3.0.
Reviewed-on: #495
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
v0.3.0 converges a zone apex NS onto its declared nameservers instead of leaving the seed placeholder, which publishes a primary pod IP. The CRD moves with the image because the operator reads the new `spec.nameservers` field.
- pin the bind-operator image to v0.3.0
- pull the CRDs from the v0.3.0 tag
Reviewed-on: #494
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The haproxy edge holds its own DMZ VIP, a second public entry point alongside
traefik-external that must be firewalled and DNS'd separately. Traefik can
front it with TLS passthrough, leaving haproxy's certs and backends untouched.
- Add a `traefik-external` Gateway: HTTP :80 plus Passthrough TLS :443.
- TLSRoute the 12 `fe_https.map` hostnames to haproxy:443; HTTPRoute 301s :80.
- Make the Service ClusterIP on 443 only, releasing 198.18.199.1.
- Drop `fe_http`, `be_letsencrypt` and `fe_http.map`; certs are DNS-01 only.
Client IP now reads as a Traefik pod — the Gateway provider cannot emit PROXY protocol to a TLSRoute backend. `sessionAffinity` goes too (it would pin Traefik pods, not clients); SRVNAME cookies keep persistence.
Reviewed-on: #492
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Stalwart was only ever a test deployment. The daemon is dead on all three
backend VMs and nothing public depends on it — `unkin.net` MX points at Google —
so the edge is proxying mail to nowhere and the tcp frontends make `defaults`
emit 20 spurious HTTP-mode warnings.
- Drop the `fe_smtp`, `fe_submission`, `fe_imap` and `fe_imaps` frontends.
- Drop the five `be_stalwart_*` backends and their map entries in `fe_http.map`/`fe_https.map`.
- Drop the now-unused 25/143/587/993 Service and container ports.
`haproxy -c` on the rendered config: exit 0, 0 warnings (was 20), 0 alerts.
Reviewed-on: #491
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Go builds on CI and laptops each rebuild the same packages from scratch. A
GOCACHEPROG backend needs an S3 bucket plus credentials before anything can
point at it, so provision those first. The bucket lives in the woodpecker
namespace because CI is the primary consumer and reads the Secret there.
- add Bucket and ObjectStoreUser for the shared Go build cache
- use default (replicated) placement rather than the ec target, since a build
cache is millions of small objects
- purge and drop the bucket and user on delete; the cache is disposable
Nothing consumes the bucket yet.
Reviewed-on: #489
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
A g10k-code job wedged in ContainerCreating on a failed CephFS mount and never reached a terminal condition, so it stayed in the CronJob active list and `concurrencyPolicy: Forbid` skipped every following minute. No Puppet code reached the estate for 6 days, and the piled-up missed slots crossed the controller 100-slot cap into `TooManyMissedTimes`. The CronJob carried no deadlines at all.
- Cap a job at `activeDeadlineSeconds: 300` on the Job spec, so a hang is failed as `DeadlineExceeded` and drops out of the active list (healthy runs take 16-18s).
- Set `startingDeadlineSeconds: 200`, bounding missed-schedule look-back to ~3 slots so the count cannot reach 100.
Reviewed-on: #490
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The 13 `X-Frame-Options DENY if acl_<host>` rules in `fe_https` have never fired:
their ACLs use `req.hdr(host)`, a request-direction fetch that is invalid in a
response ruleset, so HAProxy rejects them at config-check time. Carried over
verbatim from the Puppet LXD config during the k8s move.
- Remove the 13 dead `http-response set-header X-Frame-Options` rules.
- Remove the 15 now-orphaned `acl acl_*` definition lines.
Not switching the header on: it has never been live, and Grafana/Gitea send their
own. `haproxy -c` warnings drop 33 -> 20; the two working response headers stay.
Reviewed-on: #487
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The au-syd1 edge proxy runs on a hand-managed LXD container outside the cluster, with no HA and no shared config source.
- Add `apps/base/haproxy/`: 3 replicas behind the DMZ LoadBalancer 198.18.199.1, config from a ConfigMap, wildcard certs from reflected secrets.
- Keep source IPs via `externalTrafficPolicy: Local`; `sessionAffinity: ClientIP` stands in for the stick-table peers a Deployment cannot name.
- Drain on shutdown: `hard-stop-after 2m`, a preStop SIGUSR1 soft-stop, 150s grace.
- Bind the stats listener to 127.0.0.1 so it is port-forward only.
- Register the app in the platform project and ApplicationSet.
Reviewed-on: #485
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The acme.unkin.net zone still serves only the bind-operator seed apex: NS ns1.acme.unkin.net glued to A 10.42.6.38, a pod IP no pod holds. The parent delegates to acme-ns1.unkin.net, but public resolvers have already promoted the child NS RRset, so when the cached address expires DNS-01 fails for every unkin.net wildcard at once.
- Add apex NS acme-ns1.unkin.net., matching the parent delegation (out of zone, no glue needed).
- Point ns1.acme.unkin.net at 103.216.191.185 so resolvers holding the seeded NS name still reach the zone.
- The operator seed placeholder itself is tracked separately in bind-operator.
---------
Co-authored-by: unkin-agent <agent@unkin.net>
Reviewed-on: #486
Co-authored-by: Unkin Agent <unkin-agent@unkin.net>
Co-committed-by: Unkin Agent <unkin-agent@unkin.net>
The haproxy edge terminates TLS for hosts under `main.unkin.net` and `ceph.unkin.net`, which the single `*.unkin.net` wildcard does not cover.
- Add cert-manager Certificates for both wildcards from the `letsencrypt` ClusterIssuer.
- Reflect the minted secrets into the `haproxy` namespace.
Needs these records in the public unkin.net zone first:
`_acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.`
`_acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.`
Reviewed-on: #484
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
profiles::pki::vault and profiles::ssh::sign shell out to
/usr/local/bin/certmanager and /usr/local/bin/sshsignhost from generate()
during catalog compilation. Neither binary exists in the compiler image, so
every node using them fails to compile.
- install certmanager v0.2.0 and sshsignhost v0.1.0 onto the shared bin volume with sha256 verification
- wrap both at /usr/local/bin from a pre-default entrypoint hook, failing startup loudly if either is missing
- mount read-only Vault configs for both: kubernetes auth on k8s/au/syd1, internal CA verified rather than skipped
Reviewed-on: #482
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The deployed v0.3.3 build returns 500 from /Shows/NextUp on PostgreSQL, breaking the home screen, and lets replicas diverge: library-visibility and shared-config changes never propagate, user data (resume, played state, favourites, ratings) is overwritten between pods, and eight scheduled tasks run on every replica instead of only the scan leader. v0.4.0 carries the fixes.
- Pin cheeztv and fafflix to jellyfin-ha:v0.4.0.
No config change needed: cross-pod invalidation reuses the transcode-store Redis connection string both apps already set.
Reviewed-on: #483
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Deploy bind-operator v0.2.7. The operator seeded a fresh skeleton zone file at serial 1 over zones whose BIND journal was still on disk at a higher serial; BIND rejected the inconsistent pair (`addzone failed: out of range`) and, with a PVC per replica, the stale journal outlived restarts while every reconcile rewrote the skeleton, so it never converged. That SERVFAILed roughly 1 in 3 authoritative answers for k8s.syd1.au.unkin.net and resolvers cached the failures.
- Bumps the operator image to v0.2.7
- Bumps the CRD install pin to the v0.2.7 tag, which changes the CRDs
Expect one rolling restart of the operator Deployment as the new image lands.
Reviewed-on: #480
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The new `unkin/jellyfin-plugin-sso` fork is getting a Woodpecker pipeline, and its build step will set `serviceAccountName: jellyfin-plugin-sso`. Without the SA declared here the pipeline pods fail to schedule.
- add a bare ServiceAccount `jellyfin-plugin-sso` in the `woodpecker` namespace
- register it in the woodpecker base kustomization
The step only builds .NET code, so no Vault kube-auth role or RBAC is needed.
Reviewed-on: #479
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Authentik advertises the token endpoint with a trailing slash, but some OIDC clients (the ArgoCD iOS app) POST to /application/o/token without one; Django's APPEND_SLASH will not redirect a POST, so the token exchange gets 405 and login fails.
- Add an exact-match rule on /application/o/token to the authentik and authentik-internal HTTPRoutes.
- Rewrite it to /application/o/token/ with a URLRewrite ReplaceFullPath filter, preserving the method and the authentik-server backend.
- Leave the catch-all PathPrefix rule untouched; exact matches outrank it in Gateway API precedence.
Reviewed-on: #478
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The Authentik client for ArgoCD is now public (the iOS app can't hold
a secret), so Authentik no longer enforces client_secret on token
exchange. PKCE replaces that as the protection against
authorization-code interception.
- Add `enablePKCEAuthentication: true` to the `oidc.config` block in
`argocd-cm-patch.yaml`
- Note why PKCE is needed now that the client is public
Reviewed-on: #477
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Authentik server pods expose django_prometheus metrics on :9300, but only ldap-outpost and redis-exporter are scraped in this namespace. Add the missing per-app scrape.
- add apps/base/authentik/server-vmpodscrape.yaml selecting app.kubernetes.io/name=authentik, component=server on the metrics port
- wire it into apps/base/authentik/kustomization.yaml
Reviewed-on: #476
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The cluster preloads no statement-statistics library, so there is no per-query cost attribution in postgres and slow query paths have to be inferred from application-side metrics instead of read straight out of the database.
- preload `pg_stat_statements`
- set `pg_stat_statements.max` and `.track`, which is what makes CNPG manage the extension and create it in every database
Requires a postgres restart. Stacked on `benvin/authentik-cnpg-resources`.
Reviewed-on: #475
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The 500m CPU limit is a 50ms CFS quota per 100ms period, so the postgres pods are throttled on every burst even at ~0.01 cores average and each query pays that latency. 128MB of `shared_buffers` and a 256MB cache estimate also under-serve the planner on the joins authentik issues on its hot read paths.
- raise resources to requests `500m`/`1Gi`, limits `2`/`2Gi`
- raise `shared_buffers` to 512MB and `effective_cache_size` to 1536MB
- hold the post-incident memory headroom multiple over `shared_buffers`
Rolling restart with switchover. Stacked on `benvin/authentik-hot-standby-feedback`.
Reviewed-on: #474
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Authentik serves multi-second API reads from the CNPG hot standbys. Those reads outlive `max_standby_streaming_delay`, so recovery cancels them with `canceling statement due to conflict with recovery`, which authentik surfaces as HTTP 500 — enough to break a terraform apply mid-run.
- set `hot_standby_feedback` on so replicas report their oldest xmin to the primary and long reads stop being cancelled
- SIGHUP reload only, no restart or switchover
- retained-dead-tuple cost is negligible on a ~155MB database
Reviewed-on: #473
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Edits to the cheeztv/fafflix plugin ConfigMaps only reach the pods via the inject-plugin-config initContainer, so a config change sat inert until someone manually rolled the StatefulSet. Reloader is deployed cluster-wide with autoReloadAll disabled, so each workload has to opt in.
- annotate both StatefulSets with configmap.reloader.stakater.com/auto: "true"
Reviewed-on: #471
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The internal-CA identity.k8s.syd1.au.unkin.net host has no CA bundle mounted in the jellyfin pods, so the OIDC discovery fetch fails TLS handshake (PartialChain). Authentik's discovery response is host-relative, so the browser-facing hostname must be used, not the internal one.
- Change OidEndpoint to identity.unkin.net in fafflix plugin config
- Change OidEndpoint to identity.unkin.net in cheeztv plugin config
Reviewed-on: #470
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
v0.3.0 and v0.3.1 crash-looped on Postgres migration/reader bugs and were reverted. v0.3.2 fixes both and was validated end to end against production-baseline Postgres and valkey: full migration chain completes, all previously-500 endpoints return 200, RedisTranscodeSessionStore and scan-leader gating confirmed active.
- Bump jellyfin-ha image tag v0.2.0 -> v0.3.2 in cheeztv and fafflix statefulsets
Depends on a pre-sync duplicate-username check and fresh pg_dump of both databases.
Reviewed-on: #468
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Catalog compiles fail with `LoadError: no such file to load -- toml`: server-side functions run in the puppetserver JRuby, whose gem path is separate from the agent CRuby path this hook installs into. puppet-prod's `profiles::puppet::gems` covers both; the hook only did the agent half.
- Install toml via `puppetserver gem`, mirroring the `puppetserver_gem` resource in puppet-prod
- Note in a comment that under `set -e` a failed install takes down an already-serving compiler
Reviewed-on: #467
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
v0.3.0 still serves facts from cache and collapses non-4xx upstream rejections into a 502, so clients see stale facts and lose the real status.
- Pin the pdbmux image to v0.4.0
Reviewed-on: #466
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The deployed pin sits on v0.2.0, so pdbmux still answers malformed queries with `502 all backends failed` and resolves per-certname routes by configured backend order rather than by which backend actually owns the node.
Bump the pdbmux image pin to v0.3.0:
- Replay a unanimous upstream rejection (PuppetDB's real 400 + parse message) instead of a 502.
- Resolve per-certname routes to the node's owning backend by report freshness.
- Match backend addresses case-insensitively when redacting.
Reviewed-on: #464
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
catalog-diff compiles a host's catalog in two environments and diffs them to validate puppet-prod changes before merge, which means compiling catalogs on behalf of other nodes via POST /puppet/v4/catalog. The compilers run the image default auth.conf, where that endpoint is denied.
- add a compiler auth.conf allowing catalog-diff.main.unkin.net to POST /puppet/v4/catalog
- add a pre-default entrypoint script seeding it into conf.d, failing hard if the source is absent
- mount both onto the compiler deployment via configMapGenerator
Reviewed-on: #462
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Woodpecker's k8s backend provisions a per-pipeline workspace PVC sized by
WOODPECKER_BACKEND_K8S_VOLUME_SIZE. At 10G, large builds (e.g. .NET clone +
build output) leave too little free space for tests that hard-require free
disk headroom, failing purely on disk exhaustion.
- raise WOODPECKER_BACKEND_K8S_VOLUME_SIZE from 10G to 20Gi
Reviewed-on: #459
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
**Fix-forward companion to the #457 rollback. This is NOT the current outage fix — see below.**
## The actual outage
The UI is 503 because the Authentik application slug `artifactapi` **does not exist**. OIDC discovery 404s, so oauth2-proxy exits at startup, the Service has no ready endpoints, and Traefik answers `no available server`.
```
identity.unkin.net /application/o/artifactapi/… 404
identity.k8s.syd1.au.unkin.net /application/o/artifactapi/… 404
identity.unkin.net /application/o/repospawner/… 200
identity.unkin.net /application/o/argocd/… 200
```
Root cause is upstream in **terraform-authentik**: `ci/woodpecker/push/apply` on main HEAD `4e16401` **failed**. That apply has to succeed before any argocd-apps change can help. **This PR does not fix that.**
## What this PR does fix
#456 dropped the `combine-certs` initContainer and `OAUTH2_PROXY_PROVIDER_CA_FILES`, reasoning that `identity.unkin.net` serves a publicly trusted Let's Encrypt cert and so needs no internal CA. That holds for the browser redirect but not for oauth2-proxy's own back-channel discovery/token calls.
artifactapi is the **only one of six** oauth2-proxies in the estate without it:
| app | issuer host | `PROVIDER_CA_FILES` |
|---|---|---|
| arrproxy | identity.unkin.net | yes |
| logviewer | identity.unkin.net | yes |
| mediamark | identity.unkin.net | yes |
| repospawner | identity.unkin.net | yes |
| watchstate | identity.k8s… | yes |
| **artifactapi** | identity.unkin.net | **no** |
repospawner uses the **same public `identity.unkin.net` issuer** and still needs the internal bundle, which falsifies the removal reasoning. The existing comment on that initContainer states it plainly: *"The Authentik issuer is served behind the internal unkin.net CA."*
## Changes
- Add the `combine-certs` initContainer — byte-identical to repospawner's.
- Mount the combined bundle and set `OAUTH2_PROXY_PROVIDER_CA_FILES`.
- Reload the Deployment when `vault-ca-cert` rotates.
`vault-ca-cert` already exists in the `artifactapi` namespace (`api-deployment.yaml` uses it). `kustomize build apps/base/artifactapi` succeeds.
## Risk
Trust-only and strictly additive — it appends the internal CA to the system roots. Harmless if the back channel turns out to reach a publicly trusted endpoint after all. Expected to remove the *next* blocker, surfacing as x509, once the terraform-authentik apply lands.
## Sequencing
1. Fix and re-run terraform-authentik `push/apply` so the `artifactapi` application exists.
2. Merge this.
3. Confirm `/ui/` returns 200, then close#457 unmerged.
Only merge #457 instead if the UI must come back before step 1 can be done.
Reviewed-on: #458
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
The artifactapi web UI is open to anyone who can reach the host. Front it with Authentik SSO gated on akP-artifactapi-admin, while leaving the package-manager surfaces (/api/v1, /api/v2, /v2 docker registry, /terraform, /.well-known) untouched — dnf, containerd mirrors, buildah, terraform and CI publish steps cannot do a browser flow.
- Add the oauth2-proxy ConfigMap, Deployment, Service and VMPodScrape.
- Add the oauth-credentials VaultStaticSecret.
- Point the api-route /ui rule at oauth2-proxy and add a /oauth2 rule; the catch-all / rule still goes straight to the api Service on both listeners.
Requires terraform-authentik #34 applied and the Vault kv seed first.
Reviewed-on: #456
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
v0.2.0 ships the reports/events union, aggregate summing across backends, config-file support, and the removal of the primary/prefer ranking.
## How
- Pin the pdbmux Deployment image to `v0.2.0`.
- Leave `pdbmux-env` unchanged: `PDBMUX_LISTEN`, `PDBMUX_BACKENDS`, `PDBMUX_MERGE` are the only keys v0.2.0 reads from this ConfigMap, and `PDBMUX_PRIMARY`/`PDBMUX_PREFER` are already gone (#452).
Reviewed-on: #455
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Why: pick up v3.11.2, which moves DB migrations onto golib/pg with no behavior change.
- Bump the artifactapi API and UI image pins from v3.11.1 to v3.11.2.
Reviewed-on: #453
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
pdbmux#9 drops primary/prefer and makes configured backend order the only tie-break, so the current `old`-first list would silently reverse which PuppetDB wins.
## How
- Order `PDBMUX_BACKENDS` with `new=http://puppetdb.puppet.svc.cluster.local:8080` first and `old=http://puppetdbapi.service.consul:8080` second, URLs unchanged.
- Drop `PDBMUX_PRIMARY` and `PDBMUX_PREFER`; both already resolve to `new`, on the deployed v0.1.0 image (estate defaults) and on pdbmux main (first-backend fallback), so the rendered behaviour is unchanged today.
- Refresh the configmap and deployment comments to describe order-based precedence.
Merge this before pdbmux#9.
Reviewed-on: #452
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Why: the new vimpack repo's woodpecker pipelines set `serviceAccountName: vimpack-ci`, which does not exist in the woodpecker namespace yet.
- Add bare `vimpack-ci` ServiceAccount in `apps/base/woodpecker/` and register it in the kustomization.
Reviewed-on: #454
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
The `ghcr.io/openvoxproject/openvoxserver` image reads `OPENVOXSERVER_JAVA_ARGS` (`/etc/default/puppetserver`: `JAVA_ARGS=$OPENVOXSERVER_JAVA_ARGS`) and ships no `PUPPETSERVER_*` alias, so our heap/JMX flags have been inert since the fork switch — live masters and compilers run the image default `-Xms1024m -Xmx1024m` with no JMX.
## How
- Rename `PUPPETSERVER_JAVA_ARGS` to `OPENVOXSERVER_JAVA_ARGS` in `puppetserver-master-config`, `puppetserver-compiler-config` and `puppetserver-init-config`.
- Rename the same var on the `perms-and-dirs` init container in `deployment_puppetserver-compiler.yaml`.
- Flag values are unchanged (`-Xms1024m -Xmx3072m` plus the JMX flags); heap tuning is a separate call.
Reviewed-on: #451
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Why: pick up encapi v0.1.2, which moves DB migrations onto golib/pg with no behavior change (release pipeline green).
- Bump the encapi deployment container image from v0.1.1 to v0.1.2
Reviewed-on: #450
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
An unlimited environment timeout serves stale catalogs after code changes.
## How
- Set `OPENVOXSERVER_ENVIRONMENT_TIMEOUT: "0"` in the puppetserver master and compiler ConfigMaps.
Reviewed-on: #449
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
The new `golib` repo's Woodpecker pipeline needs a dedicated ServiceAccount to run its CI steps under.
## How
- Add bare ServiceAccount `golib-ci` in the `woodpecker` namespace and wire it into the base kustomization.
Reviewed-on: #447
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
Why: repospawner v0.1.1 stops spawned job pods from automounting the API token.
- Bump the deployment image and the matching REPOSPAWNER_IMAGE env value to v0.1.1
Reviewed-on: #446
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
repospawner v0.1.0 is built and its Vault kubernetes auth role is applied, but nothing deploys it. It turns a "I want a new repository" request into a terraform-git pull request, follows that PR to merge, and optionally activates the repo in Woodpecker, so the review gate stays where it is instead of moving into an agent's hands.
## How
- Add `apps/base/repospawner/`: namespace, ServiceAccount `repospawner`, `default` VaultAuth for VSO, and a namespaced Role/RoleBinding granting jobs create/get/list/watch/delete plus pods and pods/log reads (mirrors mediamover).
- Deployment pinned to `artifactapi.k8s.syd1.au.unkin.net/docker-internal/repospawner:v0.1.0`, one replica with the `Recreate` strategy because request state is in memory and rebuilt from Job labels; the same image reference is passed down as `REPOSPAWNER_IMAGE` so the spawned Jobs stay in step.
- Mount a projected `audience: vault` service account token at `/var/run/secrets/vault` — the app logs into Vault natively rather than through VSO — and the `repospawner-woodpecker` Secret at `/etc/repospawner/woodpecker`, optional so the server still starts and refuses `woodpecker: true` with 503 when it is absent.
- Two VaultStaticSecrets: `oauth-credentials` from `kv/kubernetes/namespace/repospawner/default/oauth-credentials` and `repospawner-woodpecker` (key `token`) from `.../default/woodpecker`, with reloader annotations on both consumers.
- oauth2-proxy front door on the watchstate/mediamark pattern, gated on `akP-repospawner-admin` via the `ak_groups` claim and re-checked by the app from `X-Forwarded-Groups`; public `repospawner.unkin.net` on the reflected wildcard and internal `repospawner.k8s.syd1.au.unkin.net` on `vault-issuer`, both routed to the oauth2 Service.
- Register the overlay in the platform ApplicationSet and AppProject, and append `repospawner` to the wildcard Certificate's two reflector namespace lists.
Depends on the terraform-authentik `repospawner` client being applied and `kv/kubernetes/namespace/repospawner/default/oauth-credentials` + `.../woodpecker` being seeded.
Reviewed-on: #445
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
arrproxy v0.6.0 applies its own schema at startup under a Postgres advisory lock and holds `/readyz` until the schema is current, so every replica is safe to roll without an external gate. The wave-1 psql `arrproxy-migrate` Job and its SQL ConfigMap now only re-run idempotent statements the app already owns — dead weight, a second source of truth for the schema, and a standing drift trap whenever the app's embedded migrations move ahead of the manifests.
## How
- Bump `arrproxy-api` and `arrproxy-ui` to `v0.6.0`.
- Delete `migrate-job.yaml` and `migrations-configmap.yaml` and drop both from the arrproxy kustomization.
- Keep the wave-0/wave-2 split: wave 2 still orders the api behind the wave-0 CNPG Cluster and VSO-synced Secrets, which is independent of the migrate Job; the stale "serve only after the wave-1 migrate Job" comment is corrected.
- Rendered diff vs `main` is exactly the two image bumps plus the `arrproxy-migrate` Job and `arrproxy-migrations` ConfigMap disappearing; `kustomize build --enable-helm apps/overlays/au-syd1/arrstack` and pre-commit both clean.
Reviewed-on: #444
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>