Compare commits

..

17 Commits

Author SHA1 Message Date
unkin-agent d4a39c44e6 add agent-observability service account and namespace-scoped RBAC
ci/woodpecker/pr/vector-test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/kubeconform Pipeline was successful
2026-09-27 00:38:38 +10:00
unkin-agent 757ae5b240 Add a VictoriaLogs cluster and point logs-ingest at it (#488)
The k8s log pipeline stores to ClickHouse via NATS+vector, while the VM estate ships journald to a separate puppet-managed VictoriaLogs cluster. Consolidating on VictoriaLogs in-cluster collapses the two paths, and the logs-ingest gateway has no clients yet so it can be repointed now, ahead of the puppet change.

- add VLCluster `logs` at v1.52.0 (2 vlinsert, 2 vlselect, 3 vlstorage, 180d retention, 250Gi each on cephrbd-fast-delete)
- cap vlstorage disk use at 220GiB per node so 180d stays time-based rather than disk-bound
- repoint the logs-ingest HTTPRoute at `vlinsert-logs:9481`
- add a VictoriaLogs Grafana datasource and install its plugin

Nothing is removed here; NATS, ClickHouse, vector, logarchiver and logviewer keep running until a follow-up drops them.

Reviewed-on: #488
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:33:58 +10:00
unkin-agent 02f877540c Add gocache serve Deployment with nginx stream sidecar (#496)
`go-cache-plugin serve` binds `127.0.0.1` only, so nothing outside the pod can reach it and laptops have no way to use the S3-backed Go cache without holding RGW credentials.

- Run `go-cache-plugin serve` against the `gocache` bucket, path-style, explicit region to skip the GetBucketLocation probe
- Add an nginx sidecar stream-proxying `:9090` to the loopback plugin port, `proxy_timeout 2h`
- Publish it on PureLB `198.18.200.11`, `externalTrafficPolicy: Local` so the client IP reaches the allow rules
- Restrict to workstation + pod CIDRs: GOCACHEPROG is unauthenticated and a poisoned entry runs in every consuming build

Merge only after `docker-internal/go-cache-plugin:v0.1.0` is published.

Reviewed-on: #496
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:32:32 +10:00
unkin-agent 148dac8ca2 Declare the acme.unkin.net nameservers (#495)
The zone was seeded with an apex `NS ns1.acme.unkin.net` glued to the primary pod IP. Both were later corrected by hand, so the live RRset and the ns1 address exist only in the zone journal -- a reseed republishes the pod IP and breaks DNS-01 for every `*.unkin.net` cert. Declaring them makes git the source of truth.

- declare the two published apex NS names
- declare the in-zone ns1 address, which a seed would otherwise glue to the pod IP

Matches what the zone serves today, so applying it changes no records. Requires bind-operator v0.3.0.

Reviewed-on: #495
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:10:13 +10:00
unkin-agent 53e5846c18 Roll bind-operator to v0.3.0 (#494)
v0.3.0 converges a zone apex NS onto its declared nameservers instead of leaving the seed placeholder, which publishes a primary pod IP. The CRD moves with the image because the operator reads the new `spec.nameservers` field.

- pin the bind-operator image to v0.3.0
- pull the CRDs from the v0.3.0 tag

Reviewed-on: #494
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:06:21 +10:00
unkin-agent 9fed5decc8 bump victoria-metrics-operator chart to 0.67.3 (#493)
The vm-system overlay pins victoria-metrics-operator chart 0.57.1 (operator v0.66.1), eight operator minors behind upstream, so the cluster runs without newer CRD fields and reconciler fixes.

- bump the victoria-metrics-operator helmChart version to 0.67.3 (operator v0.74.1)

Reviewed-on: #493
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-27 00:05:47 +10:00
unkin-agent 20077f1029 Move the haproxy edge behind the external Traefik (#492)
The haproxy edge holds its own DMZ VIP, a second public entry point alongside
traefik-external that must be firewalled and DNS'd separately. Traefik can
front it with TLS passthrough, leaving haproxy's certs and backends untouched.

- Add a `traefik-external` Gateway: HTTP :80 plus Passthrough TLS :443.
- TLSRoute the 12 `fe_https.map` hostnames to haproxy:443; HTTPRoute 301s :80.
- Make the Service ClusterIP on 443 only, releasing 198.18.199.1.
- Drop `fe_http`, `be_letsencrypt` and `fe_http.map`; certs are DNS-01 only.

Client IP now reads as a Traefik pod — the Gateway provider cannot emit PROXY protocol to a TLSRoute backend. `sessionAffinity` goes too (it would pin Traefik pods, not clients); SRVNAME cookies keep persistence.

Reviewed-on: #492
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 23:38:07 +10:00
unkin-agent 426a399f31 Drop stalwart mail proxying from the haproxy edge (#491)
Stalwart was only ever a test deployment. The daemon is dead on all three
backend VMs and nothing public depends on it — `unkin.net` MX points at Google —
so the edge is proxying mail to nowhere and the tcp frontends make `defaults`
emit 20 spurious HTTP-mode warnings.

- Drop the `fe_smtp`, `fe_submission`, `fe_imap` and `fe_imaps` frontends.
- Drop the five `be_stalwart_*` backends and their map entries in `fe_http.map`/`fe_https.map`.
- Drop the now-unused 25/143/587/993 Service and container ports.

`haproxy -c` on the rendered config: exit 0, 0 warnings (was 20), 0 alerts.

Reviewed-on: #491
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 21:25:49 +10:00
unkin-agent 5341253573 Add Ceph RGW bucket for the shared Go build cache (#489)
Go builds on CI and laptops each rebuild the same packages from scratch. A
GOCACHEPROG backend needs an S3 bucket plus credentials before anything can
point at it, so provision those first. The bucket lives in the woodpecker
namespace because CI is the primary consumer and reads the Secret there.

- add Bucket and ObjectStoreUser for the shared Go build cache
- use default (replicated) placement rather than the ec target, since a build
  cache is millions of small objects
- purge and drop the bucket and user on delete; the cache is disposable

Nothing consumes the bucket yet.

Reviewed-on: #489
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:56:51 +10:00
unkin-agent d48125d699 Add job and start deadlines to the g10k-code CronJob (#490)
A g10k-code job wedged in ContainerCreating on a failed CephFS mount and never reached a terminal condition, so it stayed in the CronJob active list and `concurrencyPolicy: Forbid` skipped every following minute. No Puppet code reached the estate for 6 days, and the piled-up missed slots crossed the controller 100-slot cap into `TooManyMissedTimes`. The CronJob carried no deadlines at all.

- Cap a job at `activeDeadlineSeconds: 300` on the Job spec, so a hang is failed as `DeadlineExceeded` and drops out of the active list (healthy runs take 16-18s).
- Set `startingDeadlineSeconds: 200`, bounding missed-schedule look-back to ~3 slots so the count cannot reach 100.

Reviewed-on: #490
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:55:32 +10:00
unkin-agent abf6bfae88 Drop dead X-Frame-Options rules from the haproxy edge (#487)
The 13 `X-Frame-Options DENY if acl_<host>` rules in `fe_https` have never fired:
their ACLs use `req.hdr(host)`, a request-direction fetch that is invalid in a
response ruleset, so HAProxy rejects them at config-check time. Carried over
verbatim from the Puppet LXD config during the k8s move.

- Remove the 13 dead `http-response set-header X-Frame-Options` rules.
- Remove the 15 now-orphaned `acl acl_*` definition lines.

Not switching the header on: it has never been live, and Grafana/Gitea send their
own. `haproxy -c` warnings drop 33 -> 20; the two working response headers stay.

Reviewed-on: #487
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 20:36:26 +10:00
unkin-agent 4190785389 Move the au-syd1 haproxy edge into Kubernetes (#485)
The au-syd1 edge proxy runs on a hand-managed LXD container outside the cluster, with no HA and no shared config source.

- Add `apps/base/haproxy/`: 3 replicas behind the DMZ LoadBalancer 198.18.199.1, config from a ConfigMap, wildcard certs from reflected secrets.
- Keep source IPs via `externalTrafficPolicy: Local`; `sessionAffinity: ClientIP` stands in for the stick-table peers a Deployment cannot name.
- Drain on shutdown: `hard-stop-after 2m`, a preStop SIGUSR1 soft-stop, 150s grace.
- Bind the stats listener to 127.0.0.1 so it is port-forward only.
- Register the app in the platform project and ApplicationSet.

Reviewed-on: #485
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 18:22:57 +10:00
unkin-agent 812a9a2f2b Publish real delegation records for acme.unkin.net (#486)
The acme.unkin.net zone still serves only the bind-operator seed apex: NS ns1.acme.unkin.net glued to A 10.42.6.38, a pod IP no pod holds. The parent delegates to acme-ns1.unkin.net, but public resolvers have already promoted the child NS RRset, so when the cached address expires DNS-01 fails for every unkin.net wildcard at once.

- Add apex NS acme-ns1.unkin.net., matching the parent delegation (out of zone, no glue needed).
- Point ns1.acme.unkin.net at 103.216.191.185 so resolvers holding the seeded NS name still reach the zone.
- The operator seed placeholder itself is tracked separately in bind-operator.

---------

Co-authored-by: unkin-agent <agent@unkin.net>
Reviewed-on: #486
Co-authored-by: Unkin Agent <unkin-agent@unkin.net>
Co-committed-by: Unkin Agent <unkin-agent@unkin.net>
2026-09-26 16:42:02 +10:00
unkin-agent f37749523d Add *.main and *.ceph wildcard certificates for haproxy (#484)
The haproxy edge terminates TLS for hosts under `main.unkin.net` and `ceph.unkin.net`, which the single `*.unkin.net` wildcard does not cover.

- Add cert-manager Certificates for both wildcards from the `letsencrypt` ClusterIssuer.
- Reflect the minted secrets into the `haproxy` namespace.

Needs these records in the public unkin.net zone first:
`_acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.`
`_acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.`

Reviewed-on: #484
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-26 16:38:31 +10:00
unkin-agent fe51aa07be Give the puppetserver compilers the Vault cert helpers (#482)
profiles::pki::vault and profiles::ssh::sign shell out to
/usr/local/bin/certmanager and /usr/local/bin/sshsignhost from generate()
during catalog compilation. Neither binary exists in the compiler image, so
every node using them fails to compile.

- install certmanager v0.2.0 and sshsignhost v0.1.0 onto the shared bin volume with sha256 verification
- wrap both at /usr/local/bin from a pre-default entrypoint hook, failing startup loudly if either is missing
- mount read-only Vault configs for both: kubernetes auth on k8s/au/syd1, internal CA verified rather than skipped

Reviewed-on: #482
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-24 21:15:09 +10:00
unkin-agent cdaab736b5 Bump jellyfin-ha to v0.4.0 (#483)
The deployed v0.3.3 build returns 500 from /Shows/NextUp on PostgreSQL, breaking the home screen, and lets replicas diverge: library-visibility and shared-config changes never propagate, user data (resume, played state, favourites, ratings) is overwritten between pods, and eight scheduled tasks run on every replica instead of only the scan leader. v0.4.0 carries the fixes.

- Pin cheeztv and fafflix to jellyfin-ha:v0.4.0.

No config change needed: cross-pod invalidation reuses the transcode-store Redis connection string both apps already set.

Reviewed-on: #483
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-22 22:57:06 +10:00
unkin-agent b31517e6d9 Merge pull request #481 from benvin/jellyfin-sso-valkey-state
Roll jellyfin-ha to v0.3.3 and drop Service session affinity

Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
2026-09-20 14:02:48 +10:00
47 changed files with 1228 additions and 17 deletions
@@ -7,4 +7,5 @@ resources:
- cluster.yaml
- tsigkey.yaml
- zones.yaml
- records.yaml
- agent-dns-rolebinding.yaml
+36
View File
@@ -0,0 +1,36 @@
# Authoritative delegation records for acme.unkin.net. Without these the zone
# only holds the operator's seed apex (NS ns1.acme.unkin.net glued to the
# primary pod IP), which is unroutable off-cluster and goes stale on
# reschedule. DNSRecords must live in the same namespace as their BindZone.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: acme-apex-ns
namespace: bind-external
spec:
zoneRef: acme-unkin-net
# "@" is the zone apex.
name: "@"
type: NS
ttl: 3600
values:
# Matches the parent delegation in Google Cloud DNS. Out of zone, so the
# child needs no glue of its own.
- acme-ns1.unkin.net.
---
apiVersion: bind.unkin.net/v1alpha1
kind: DNSRecord
metadata:
name: acme-ns1-a
namespace: bind-external
spec:
zoneRef: acme-unkin-net
name: ns1
type: A
ttl: 3600
values:
# Public address of this cluster's external BIND, same target as
# acme-ns1.unkin.net. Resolvers that cached the seeded ns1.acme.unkin.net
# NS name must still reach the zone.
- 103.216.191.185
+11
View File
@@ -17,3 +17,14 @@ spec:
updateKeyRef: certmanager
allowTransfer:
- key certmanager
# Published apex NS. acme-ns1 is what the parent delegates to and glues; ns1 is
# in-zone, so its address is declared below or a reseed would glue it to the
# primary pod IP.
nameservers:
- acme-ns1.unkin.net.
- ns1.acme.unkin.net.
records:
- name: ns1
type: A
ttl: 3600
values: ["103.216.191.185"]
+1 -1
View File
@@ -21,7 +21,7 @@ spec:
runAsNonRoot: true
containers:
- name: operator
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/bind-operator:v0.2.7
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/bind-operator:v0.3.0
args:
- --metrics-bind-address=:8080
- --health-probe-bind-address=:8081
+1 -1
View File
@@ -6,7 +6,7 @@ resources:
- namespace.yaml
# CRDs are pulled from the bind-operator repo at the matching tag rather than
# vendored here, so they never drift from the operator.
- https://git.unkin.net/unkin/bind-operator/raw/tag/v0.2.7/config/crd/install.yaml
- https://git.unkin.net/unkin/bind-operator/raw/tag/v0.3.0/config/crd/install.yaml
- rbac.yaml
- agent-dns-rbac.yaml
- deployment.yaml
@@ -0,0 +1,26 @@
---
# Let's Encrypt *.ceph.unkin.net wildcard for the haproxy edge (ceph dashboard).
# DNS-01 needs the delegated _acme-challenge.ceph.unkin.net CNAME in the public
# unkin.net zone.
# _acme-challenge.ceph.unkin.net. CNAME _acme-challenge.ceph.acme.unkin.net.
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: wildcard-ceph-unkin-net
namespace: cert-manager
spec:
secretName: wildcard-ceph-unkin-net-tls
secretTemplate:
annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "haproxy"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "haproxy"
privateKey:
size: 4096
dnsNames:
- "*.ceph.unkin.net"
issuerRef:
name: letsencrypt
kind: ClusterIssuer
group: cert-manager.io
@@ -0,0 +1,26 @@
---
# Let's Encrypt *.main.unkin.net wildcard for the haproxy edge (pve, arr stack,
# jellyfin, stalwart webadmin/autoconfig). DNS-01 needs the delegated
# _acme-challenge.main.unkin.net CNAME in the public unkin.net zone.
# _acme-challenge.main.unkin.net. CNAME _acme-challenge.main.acme.unkin.net.
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: wildcard-main-unkin-net
namespace: cert-manager
spec:
secretName: wildcard-main-unkin-net-tls
secretTemplate:
annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "haproxy"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "haproxy"
privateKey:
size: 4096
dnsNames:
- "*.main.unkin.net"
issuerRef:
name: letsencrypt
kind: ClusterIssuer
group: cert-manager.io
@@ -14,9 +14,9 @@ spec:
secretTemplate:
annotations:
reflector.v1.k8s.emberstack.com/reflection-allowed: "true"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner"
reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner,haproxy"
reflector.v1.k8s.emberstack.com/reflection-auto-enabled: "true"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner"
reflector.v1.k8s.emberstack.com/reflection-auto-namespaces: "cheeztv,arrstack,authentik,gitea,watchstate,mediamark,repospawner,haproxy"
privateKey:
size: 4096
dnsNames:
@@ -12,3 +12,5 @@ resources:
- clusterissuer_letsencrypt.yaml
- clusterissuer_letsencrypt-staging.yaml
- certificate_wildcard-unkin-net.yaml
- certificate_wildcard-main-unkin-net.yaml
- certificate_wildcard-ceph-unkin-net.yaml
+1 -1
View File
@@ -164,7 +164,7 @@ spec:
readOnly: true
containers:
- name: cheeztv
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.3.3
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.4.0
imagePullPolicy: IfNotPresent
ports:
- name: http
+1 -1
View File
@@ -164,7 +164,7 @@ spec:
readOnly: true
containers:
- name: fafflix
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.3.3
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/jellyfin-ha:v0.4.0
imagePullPolicy: IfNotPresent
ports:
- name: http
+19
View File
@@ -20,3 +20,22 @@ spec:
jsonData:
timeInterval: "15s"
httpMethod: "POST"
---
apiVersion: grafana.integreatly.org/v1beta1
kind: GrafanaDatasource
metadata:
name: victorialogs
namespace: grafana
spec:
instanceSelector:
matchLabels:
dashboards: "grafana"
plugins:
- name: victoriametrics-logs-datasource
version: 0.32.0
datasource:
name: "VictoriaLogs"
type: "victoriametrics-logs-datasource"
uid: "victorialogs"
access: "proxy"
url: "http://vlselect-logs.logging.svc.cluster.local:9471"
+274
View File
@@ -0,0 +1,274 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: haproxy-config
namespace: haproxy
data:
certificate.list: |
# First entry is the default cert for non-matching SNI.
/etc/haproxy/certs/unkin-net/tls.crt
/etc/haproxy/certs/main-unkin-net/tls.crt
/etc/haproxy/certs/ceph-unkin-net/tls.crt
fe_https.map: |
sonarr.main.unkin.net be_sonarr
radarr.main.unkin.net be_radarr
lidarr.main.unkin.net be_lidarr
readarr.main.unkin.net be_readarr
prowlarr.main.unkin.net be_prowlarr
nzbget.main.unkin.net be_nzbget
jellyfin.main.unkin.net be_jellyfin
fafflix.unkin.net be_jellyfin
git.unkin.net be_gitea
grafana.unkin.net be_grafana
dashboard.ceph.unkin.net be_ceph_dashboard
auth.unkin.net be_k8s_kanidm
haproxy.cfg: |
global
log stdout format raw local0
log stdout format raw local1 notice
maxconn 4000
hard-stop-after 2m
ssl-default-bind-ciphers EECDH+AESGCM:EDH+AESGCM:AES256+EECDH:AES256+EDH
ssl-default-bind-options ssl-min-ver TLSv1.2 ssl-max-ver TLSv1.3
ssl-default-server-ciphers kEECDH+aRSA+AES:kRSA+AES:+AES256:RC4-SHA:!kEDH:!LOW:!EXP:!MD5:!aNULL:!eNULL
ssl-default-server-options no-sslv3
stats timeout 30s
stats socket /var/lib/haproxy/stats
stats socket /var/lib/haproxy/admin.sock mode 660 level admin
tune.ssl.default-dh-param 2048
defaults
log global
maxconn 5000
mode http
option httplog
option dontlognull
option http-server-close
option forwardfor except 127.0.0.0/8
option redispatch
retries 3
stats enable
timeout http-request 10s
timeout queue 1m
timeout connect 10s
timeout client 5m
timeout server 5m
timeout http-keep-alive 10s
timeout check 10s
frontend fe_https
bind 0.0.0.0:443 ssl crt-list /usr/local/etc/haproxy/certificate.list ciphers EECDH+AESGCM:EDH+AESGCM:AES256+EECDH:AES256+EDH force-tlsv12
mode http
description Global HTTPS Frontend
http-request set-header X-Forwarded-Proto https
http-request set-header X-Real-IP %[src]
http-response set-header X-Content-Type-Options nosniff
http-response set-header X-XSS-Protection 1;mode=block
use_backend %[req.hdr(host),lower,map(/usr/local/etc/haproxy/fe_https.map,be_default)]
frontend fe_metrics
bind 0.0.0.0:8405
mode http
description Metrics Frontend
http-request set-header X-Forwarded-Proto https
http-request set-header X-Real-IP %[src]
http-request use-service prometheus-exporter if { path /metrics }
backend be_ceph_dashboard
description Backend for Ceph Dashboard from Mgr instances
balance roundrobin
cookie SRVNAME insert indirect nocache
http-check expect status 200
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 9443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
stick-table type ip size 200k expire 30m
server prodnxsr0009 198.18.23.9:9443 check cookie prodnxsr0009 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0010 198.18.23.10:9443 check cookie prodnxsr0010 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0011 198.18.23.11:9443 check cookie prodnxsr0011 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0012 198.18.23.12:9443 check cookie prodnxsr0012 fall 2 inter 2s rise 3 ssl verify none
server prodnxsr0013 198.18.23.13:9443 check cookie prodnxsr0013 fall 2 inter 2s rise 3 ssl verify none
backend be_default
description Backend for unmatched HTTP traffic
balance roundrobin
cookie SRVNAME insert
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
option httpchk GET /
option forwardfor
backend be_gitea
description Backend for gitea cluster
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
stick on src
stick-table type ip size 200k expire 30m
server ausyd1nxvm2080 198.18.26.18:443 check cookie ausyd1nxvm2080 fall 2 inter 2s rise 3 ssl verify none
server ausyd1nxvm2081 198.18.27.117:443 check cookie ausyd1nxvm2081 fall 2 inter 2s rise 3 ssl verify none
server ausyd1nxvm2082 198.18.28.71:443 check cookie ausyd1nxvm2082 fall 2 inter 2s rise 3 ssl verify none
backend be_grafana
description Backend for grafana nodes
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
stick on src
stick-table type ip size 200k expire 30m
server ausyd1nxvm2015 198.18.27.2:443 check cookie ausyd1nxvm2015 fall 2 inter 2s rise 3 ssl verify none
server ausyd1nxvm2016 198.18.28.189:443 check cookie ausyd1nxvm2016 fall 2 inter 2s rise 3 ssl verify none
backend be_jellyfin
description Backend for au-syd1 jellyfin
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2051 198.18.25.164:443 check cookie ausyd1nxvm2051 fall 2 inter 2s rise 3 ssl verify none
backend be_k8s_kanidm
description Backend for Kanidm (auth.unkin.net via Kubernetes internal Traefik)
balance roundrobin
http-reuse always
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
redirect scheme https if !{ ssl_fc }
option httpchk
option forwardfor
option http-keep-alive
option prefer-last-server
http-check connect ssl sni auth.unkin.net
http-check send meth GET uri /status ver HTTP/1.1 hdr Host auth.unkin.net
http-check expect status 200
server k8s-traefik-internal 198.18.200.4:443 ssl verify none check inter 2s rise 3 fall 2 sni str(auth.unkin.net)
backend be_lidarr
description Backend for au-syd1 lidarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2048 198.18.28.165:443 check cookie ausyd1nxvm2048 fall 2 inter 2s rise 3 ssl verify none
backend be_nzbget
description Backend for au-syd1 nzbget
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2045 198.18.25.44:443 check cookie ausyd1nxvm2045 fall 2 inter 2s rise 3 ssl verify none
backend be_prowlarr
description Backend for au-syd1 prowlarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2050 198.18.25.66:443 check cookie ausyd1nxvm2050 fall 2 inter 2s rise 3 ssl verify none
backend be_radarr
description Backend for au-syd1 radarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2047 198.18.27.131:443 check cookie ausyd1nxvm2047 fall 2 inter 2s rise 3 ssl verify none
backend be_readarr
description Backend for au-syd1 readarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2049 198.18.29.32:443 check cookie ausyd1nxvm2049 fall 2 inter 2s rise 3 ssl verify none
backend be_sonarr
description Backend for au-syd1 sonarr
balance roundrobin
cookie SRVNAME insert indirect nocache
http-request set-header X-Forwarded-Port %[dst_port]
http-request add-header X-Forwarded-Proto https if { dst_port 443 }
http-reuse always
option httpchk GET /consul/health
option forwardfor
option http-keep-alive
option prefer-last-server
redirect scheme https if !{ ssl_fc }
server ausyd1nxvm2046 198.18.26.161:443 check cookie ausyd1nxvm2046 fall 2 inter 2s rise 3 ssl verify none
# The `peers au-syd1-prod` section is dropped: peer names must be static and a
# Deployment cannot provide them. Behind the external Traefik's TLS
# passthrough `src` is a Traefik pod, so X-Real-IP, forwardfor and the
# `stick on src` tables all key on that; the SRVNAME cookie carries real
# session persistence. Traefik cannot emit PROXY protocol to a TLSRoute
# backend, so there is nothing to bind `accept-proxy` to.
listen health
bind 0.0.0.0:8404
mode http
monitor-uri /healthz
listen stats
bind 127.0.0.1:9090
mode http
stats uri /
stats auth admin:admin
+148
View File
@@ -0,0 +1,148 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: haproxy
namespace: haproxy
annotations:
reloader.stakater.com/auto: "true"
spec:
replicas: 3
selector:
matchLabels:
app: haproxy
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app: haproxy
spec:
automountServiceAccountToken: false
terminationGracePeriodSeconds: 150
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
app: haproxy
topologyKey: kubernetes.io/hostname
securityContext:
runAsNonRoot: true
runAsUser: 99
runAsGroup: 99
seccompProfile:
type: RuntimeDefault
containers:
- name: haproxy
image: haproxy:3.2.24-alpine
imagePullPolicy: IfNotPresent
command:
- haproxy
- -W
- -db
- -f
- /usr/local/etc/haproxy/haproxy.cfg
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: [ALL]
# fe_https binds the privileged port 443 as uid 99, and the
# dst_port ACLs need the real port.
add: [NET_BIND_SERVICE]
ports:
- name: https
containerPort: 443
protocol: TCP
- name: health
containerPort: 8404
protocol: TCP
- name: metrics
containerPort: 8405
protocol: TCP
- name: stats
containerPort: 9090
protocol: TCP
lifecycle:
preStop:
exec:
# SIGUSR1 to the master soft-stops the workers; hard-stop-after
# caps the drain. Wait so kubelet holds SIGTERM until it is done.
command:
- /bin/sh
- -c
- kill -s USR1 1; while kill -0 1 2>/dev/null; do sleep 1; done
livenessProbe:
httpGet:
path: /healthz
port: health
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /healthz
port: health
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 5
failureThreshold: 3
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: 2
memory: 1Gi
volumeMounts:
- name: config
mountPath: /usr/local/etc/haproxy
readOnly: true
- name: cert-unkin-net
mountPath: /etc/haproxy/certs/unkin-net
readOnly: true
- name: cert-main-unkin-net
mountPath: /etc/haproxy/certs/main-unkin-net
readOnly: true
- name: cert-ceph-unkin-net
mountPath: /etc/haproxy/certs/ceph-unkin-net
readOnly: true
- name: run
mountPath: /var/lib/haproxy
volumes:
- name: config
configMap:
name: haproxy-config
# ssl-load-extra-files loads <crtfile>.key by default, so the key is
# projected next to the cert as tls.crt.key.
- name: cert-unkin-net
secret:
secretName: wildcard-unkin-net-tls
items:
- key: tls.crt
path: tls.crt
- key: tls.key
path: tls.crt.key
- name: cert-main-unkin-net
secret:
secretName: wildcard-main-unkin-net-tls
items:
- key: tls.crt
path: tls.crt
- key: tls.key
path: tls.crt.key
- name: cert-ceph-unkin-net
secret:
secretName: wildcard-ceph-unkin-net-tls
items:
- key: tls.crt
path: tls.crt
- key: tls.key
path: tls.crt.key
- name: run
emptyDir: {}
restartPolicy: Always
+31
View File
@@ -0,0 +1,31 @@
---
# External (DMZ) front for the haproxy edge on the traefik-external LB VIP
# 198.18.199.0. The :443 listener is TLS Passthrough: haproxy owns the three
# wildcard certs and terminates behind Traefik, so there are no certificateRefs
# here. Listener hostnames are deliberately unset and the routes carry the
# explicit hostname list instead; allowedRoutes Same keeps other namespaces off
# these listeners.
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: haproxy
namespace: haproxy
labels:
traefik.io/instance: external
spec:
gatewayClassName: traefik-external
listeners:
- name: http
port: 80
protocol: HTTP
allowedRoutes:
namespaces:
from: Same
- name: https-passthrough
port: 443
protocol: TLS
tls:
mode: Passthrough
allowedRoutes:
namespaces:
from: Same
+37
View File
@@ -0,0 +1,37 @@
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: haproxy-http-redirect
namespace: haproxy
labels:
app: haproxy
spec:
hostnames:
- sonarr.main.unkin.net
- radarr.main.unkin.net
- lidarr.main.unkin.net
- readarr.main.unkin.net
- prowlarr.main.unkin.net
- nzbget.main.unkin.net
- jellyfin.main.unkin.net
- fafflix.unkin.net
- git.unkin.net
- grafana.unkin.net
- dashboard.ceph.unkin.net
- auth.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: haproxy
sectionName: http
rules:
- filters:
- type: RequestRedirect
requestRedirect:
scheme: https
statusCode: 301
matches:
- path:
type: PathPrefix
value: /
+15
View File
@@ -0,0 +1,15 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- configmap.yaml
- deployment.yaml
- service.yaml
- gateway.yaml
- tlsroute.yaml
- httproute.yaml
- pdb.yaml
- vpa.yaml
- vmpodscrape.yaml
+5
View File
@@ -0,0 +1,5 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: haproxy
+11
View File
@@ -0,0 +1,11 @@
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: haproxy
namespace: haproxy
spec:
maxUnavailable: 1
selector:
matchLabels:
app: haproxy
+19
View File
@@ -0,0 +1,19 @@
---
apiVersion: v1
kind: Service
metadata:
name: haproxy
namespace: haproxy
spec:
type: ClusterIP
# Reached only by the external Traefik's TLS-passthrough TLSRoute, so the
# peer address here is a Traefik pod, not the client. sessionAffinity is
# deliberately absent: keyed on ClientIP it would pin whole Traefik pods,
# not clients. Backend persistence rests on the per-backend SRVNAME cookie.
selector:
app: haproxy
ports:
- name: https
port: 443
protocol: TCP
targetPort: https
+34
View File
@@ -0,0 +1,34 @@
---
apiVersion: gateway.networking.k8s.io/v1
kind: TLSRoute
metadata:
name: haproxy
namespace: haproxy
labels:
app: haproxy
spec:
hostnames:
- sonarr.main.unkin.net
- radarr.main.unkin.net
- lidarr.main.unkin.net
- readarr.main.unkin.net
- prowlarr.main.unkin.net
- nzbget.main.unkin.net
- jellyfin.main.unkin.net
- fafflix.unkin.net
- git.unkin.net
- grafana.unkin.net
- dashboard.ceph.unkin.net
- auth.unkin.net
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: haproxy
sectionName: https-passthrough
rules:
- backendRefs:
- group: ""
kind: Service
name: haproxy
port: 443
weight: 1
+13
View File
@@ -0,0 +1,13 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: haproxy
namespace: haproxy
spec:
selector:
matchLabels:
app: haproxy
podMetricsEndpoints:
- port: metrics
path: /metrics
+13
View File
@@ -0,0 +1,13 @@
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: haproxy-vpa
namespace: haproxy
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: haproxy
updatePolicy:
updateMode: "Off"
@@ -0,0 +1,16 @@
---
# Confines the agent-observability service account (in vm-system) to the
# agent-observability ClusterRole within this namespace.
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: agent-observability
namespace: logging
subjects:
- kind: ServiceAccount
name: agent-observability
namespace: vm-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: agent-observability
+3 -6
View File
@@ -1,16 +1,13 @@
---
# Log ingestion endpoint for puppet-managed VMs (and any non-k8s client).
# Reuses the internal Traefik gateway + cert-manager + external-dns pattern so
# VMs reach the Vector aggregator's HTTP source over TLS at a DNS name they can
# resolve. The puppet-side Vector rollout ships NDJSON to
# https://logs-ingest.k8s.syd1.au.unkin.net/ (a later task).
# Log ingestion endpoint for puppet-managed VMs (and any non-k8s client):
# fronts the VLCluster vlinsert service over TLS at a name VMs can resolve.
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: logs-ingest
namespace: logging
labels:
app.kubernetes.io/name: vector-aggregator
app.kubernetes.io/name: victorialogs
app.kubernetes.io/component: ingest
traefik.io/instance: internal
annotations:
+4 -4
View File
@@ -5,7 +5,7 @@ metadata:
name: logs-ingest-http-redirect
namespace: logging
labels:
app.kubernetes.io/name: vector-aggregator
app.kubernetes.io/name: victorialogs
app.kubernetes.io/component: ingest
spec:
hostnames:
@@ -32,7 +32,7 @@ metadata:
name: logs-ingest
namespace: logging
labels:
app.kubernetes.io/name: vector-aggregator
app.kubernetes.io/name: victorialogs
app.kubernetes.io/component: ingest
spec:
hostnames:
@@ -46,8 +46,8 @@ spec:
- backendRefs:
- group: ""
kind: Service
name: vector-vm-ingest
port: 8080
name: vlinsert-logs
port: 9481
weight: 1
matches:
- path:
+2
View File
@@ -10,12 +10,14 @@ resources:
- job_clickhouse-schema.yaml
- nats-bootstrap-job.yaml
- cephrgw.yaml
- vlcluster.yaml
- gateway.yaml
- httproute.yaml
- serviceaccount_logarchiver.yaml
- configmap_logarchiver.yaml
- deployment_logarchiver.yaml
- logviewer
- agent-observability-rolebinding.yaml
# Vector pipelines are the single source of truth (also validated by
# `vector test` in CI). Mounted into each tier via `existingConfigMaps`.
+47
View File
@@ -0,0 +1,47 @@
---
apiVersion: operator.victoriametrics.com/v1
kind: VLCluster
metadata:
name: logs
namespace: logging
spec:
clusterVersion: v1.52.0
vlinsert:
replicaCount: 2
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
cpu: "2"
memory: 4Gi
vlselect:
replicaCount: 2
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
cpu: "2"
memory: 4Gi
vlstorage:
replicaCount: 3
retentionPeriod: 180d
# ~3 GiB/day measured; 220GiB/node cap keeps 180d time-based, not disk-bound
retentionMaxDiskSpaceUsageBytes: 220GiB
storage:
volumeClaimTemplate:
spec:
accessModes:
- ReadWriteOnce
storageClassName: cephrbd-fast-delete
resources:
requests:
storage: 250Gi
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "4"
memory: 8Gi
@@ -0,0 +1,16 @@
---
# Confines the agent-observability service account (in vm-system) to the
# agent-observability ClusterRole within this namespace.
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: agent-observability
namespace: observability
subjects:
- kind: ServiceAccount
name: agent-observability
namespace: vm-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: agent-observability
@@ -13,3 +13,4 @@ resources:
- httproute.yaml
- vmpodscrape-cnpg.yaml
- vmpodscrape-traefik.yaml
- agent-observability-rolebinding.yaml
+2
View File
@@ -11,11 +11,13 @@ metadata:
namespace: puppet
spec:
schedule: "*/1 * * * *"
startingDeadlineSeconds: 200
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
activeDeadlineSeconds: 300
template:
metadata:
labels:
@@ -105,6 +105,17 @@ spec:
- mountPath: /docker-custom-entrypoint.d/pre-default/10-auth-conf.sh
name: compiler-auth-conf-seed
subPath: 10-auth-conf.sh
- mountPath: /docker-custom-entrypoint.d/pre-default/20-vault-helpers.sh
name: compiler-vault-helpers-seed
subPath: 20-vault-helpers.sh
- mountPath: /opt/certmanager/config.yaml
name: certmanager-config
subPath: certmanager.yaml
readOnly: true
- mountPath: /opt/sshsignhost/config.yaml
name: sshsignhost-config
subPath: sshsignhost.yaml
readOnly: true
initContainers:
- name: copy-configmaps
image: busybox:1.35
@@ -202,7 +213,38 @@ spec:
echo "$EXPECTED encapic" | sha256sum -c -
install -m 0755 encapic /opt/bin/encapic
# Puppet shells out to these two from generate() during catalog
# compilation: profiles::pki::vault runs certmanager and
# profiles::ssh::sign runs sshsignhost.
install_release() {
name=$1
version=$2
asset="$name-linux-amd64"
base="https://git.unkin.net/unkin/$name/releases/download/$version"
curl -fsSL -o "$name" "$base/$asset"
curl -fsSL -o "$name.checksums" "$base/checksums.txt"
# checksums.txt covers every release asset; pick the line for the
# one we downloaded and verify it under our local filename.
expected=$(awk -v a="$asset" '$NF == a || $NF == "*"a {print $1}' "$name.checksums")
if [ -z "$expected" ]; then
echo "no checksum for $asset in $version checksums.txt" >&2
exit 1
fi
echo "$expected $name" | sha256sum -c -
install -m 0755 "$name" "/opt/bin/$name"
}
install_release certmanager v0.2.0
install_release sshsignhost v0.1.0
echo "Shared binaries setup completed"
resources:
limits:
cpu: 300m
memory: 256Mi
requests:
cpu: 100m
memory: 64Mi
volumeMounts:
- mountPath: /opt/bin/
name: puppet-shared-bins
@@ -247,5 +289,15 @@ spec:
configMap:
name: compiler-auth-conf-seed
defaultMode: 0755
- name: compiler-vault-helpers-seed
configMap:
name: compiler-vault-helpers-seed
defaultMode: 0755
- name: certmanager-config
configMap:
name: certmanager-config
- name: sshsignhost-config
configMap:
name: sshsignhost-config
strategy:
type: RollingUpdate
+15
View File
@@ -64,6 +64,21 @@ configMapGenerator:
- resources/compiler/10-auth-conf.sh
options:
disableNameSuffixHash: true
- name: compiler-vault-helpers-seed
files:
- resources/compiler/20-vault-helpers.sh
options:
disableNameSuffixHash: true
- name: certmanager-config
files:
- resources/compiler/certmanager.yaml
options:
disableNameSuffixHash: true
- name: sshsignhost-config
files:
- resources/compiler/sshsignhost.yaml
options:
disableNameSuffixHash: true
- name: additional-ruby-gems
files:
- resources/additional-ruby-gems.sh
+29
View File
@@ -0,0 +1,29 @@
#!/bin/bash
set -euo pipefail
BIN_DIR=/opt/bin
CA=/opt/vault-ca-cert.crt
if [ ! -s "$CA" ]; then
echo "FATAL: $CA missing or empty; certmanager and sshsignhost cannot verify Vault" >&2
exit 1
fi
# profiles::pki::vault and profiles::ssh::sign shell out to fixed /usr/local/bin
# paths from generate(); the binaries ship on the shared PVC, and /usr/local/bin
# lives in the image. Wrappers rather than symlinks because neither binary reads
# a CA path from its config: SSL_CERT_FILE scopes the internal CA to these two
# processes instead of the puppetserver JVM's own trust store.
for bin in certmanager sshsignhost; do
if [ ! -x "$BIN_DIR/$bin" ]; then
echo "FATAL: $BIN_DIR/$bin missing; generate() would abort every catalog compile" >&2
exit 1
fi
cat > "/usr/local/bin/$bin" <<WRAPPER
#!/bin/sh
SSL_CERT_FILE=$CA
export SSL_CERT_FILE
exec $BIN_DIR/$bin "\$@"
WRAPPER
chmod 0755 "/usr/local/bin/$bin"
done
@@ -0,0 +1,12 @@
---
vault:
addr: https://vault.service.consul:8200
auth_method: kubernetes
k8s_mount: k8s/au/syd1
k8s_role: puppet_certmanager
jwt_path: /var/run/secrets/kubernetes.io/serviceaccount/token
mount_point: pki_int
role_name: servers_default
output_path: /tmp/certmanager
tls_skip_verify: false
timeout: 30s
@@ -0,0 +1,11 @@
---
vault:
addr: https://vault.service.consul:8200
auth_method: kubernetes
k8s_mount: k8s/au/syd1
k8s_role: puppet_sshsigner
jwt_path: /var/run/secrets/kubernetes.io/serviceaccount/token
mount_point: sshca
role_name: signhost
tls_skip_verify: false
timeout: 30s
@@ -0,0 +1,50 @@
---
# Static service account that Vault's kubernetes secret engine mints scoped
# tokens for (agent-observability role). RBAC is confined to the metrics and
# logging namespaces via the per-namespace RoleBindings, not a
# ClusterRoleBinding. Workloads and VictoriaMetrics CRs are deliberately
# patch/update only: deleting a VMCluster or VLCluster destroys data.
apiVersion: v1
kind: ServiceAccount
metadata:
name: agent-observability
namespace: vm-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: agent-observability
rules:
- apiGroups: ["operator.victoriametrics.com"]
resources: ["*"]
verbs: ["get", "list", "watch", "create", "patch", "update"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets", "daemonsets"]
verbs: ["get", "list", "watch", "patch", "update"]
# delete permits a rolling restart without granting workload deletion.
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch", "delete"]
- apiGroups: [""]
resources: ["pods/log"]
verbs: ["get"]
- apiGroups: [""]
resources: ["services", "configmaps", "endpoints", "events"]
verbs: ["get", "list", "watch"]
- apiGroups: ["gateway.networking.k8s.io"]
resources: ["gateways", "httproutes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: agent-observability
namespace: vm-system
subjects:
- kind: ServiceAccount
name: agent-observability
namespace: vm-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: agent-observability
+1
View File
@@ -4,3 +4,4 @@ kind: Kustomization
resources:
- namespace.yaml
- agent-observability-rbac.yaml
@@ -0,0 +1,39 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: gocache-nginx
namespace: woodpecker
labels:
app.kubernetes.io/name: gocache
app.kubernetes.io/component: proxy
data:
nginx.conf: |
worker_processes auto;
error_log /dev/stderr warn;
pid /tmp/nginx.pid;
events {
worker_connections 512;
}
# GOCACHEPROG is a raw byte stream, not HTTP, so this must be stream{} not http{}.
stream {
server {
listen 9090;
# The protocol has no authentication: anyone who can reach this port can
# write cache entries, which become code in every build that reads them.
# Loopback is the kubectl port-forward fallback; in a pod netns it is
# only these two containers.
allow 127.0.0.1/32;
allow 10.10.12.200/32;
allow 10.42.0.0/16;
deny all;
# A connect session lasts the whole build; the 10m default cuts long builds off.
proxy_timeout 2h;
proxy_connect_timeout 5s;
proxy_pass 127.0.0.1:9080;
}
}
@@ -0,0 +1,131 @@
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: gocache
namespace: woodpecker
annotations:
configmap.reloader.stakater.com/reload: "gocache-nginx"
secret.reloader.stakater.com/reload: "gocache-s3,vault-ca-cert"
labels:
app.kubernetes.io/name: gocache
app.kubernetes.io/component: cache
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/name: gocache
template:
metadata:
labels:
app.kubernetes.io/name: gocache
app.kubernetes.io/component: cache
spec:
serviceAccountName: default
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
fsGroup: 65532
seccompProfile:
type: RuntimeDefault
containers:
- name: go-cache-plugin
image: artifactapi.k8s.syd1.au.unkin.net/docker-internal/go-cache-plugin:v0.1.0
imagePullPolicy: IfNotPresent
# Root flags must precede the subcommand; only --plugin belongs to serve.
args:
- --cache-dir=/var/cache/gocache
- --bucket=gocache
# Explicit region skips the GetBucketLocation probe, which RGW handles poorly.
- --region=us-east-1
- --s3-endpoint-url=https://s3.ceph.unkin.net
- --s3-path-style
- serve
- --plugin=9080
env:
- name: AWS_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: gocache-s3
key: AWS_ACCESS_KEY_ID
- name: AWS_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: gocache-s3
key: AWS_SECRET_ACCESS_KEY
# s3.ceph.unkin.net is served by the estate CA, not a public root.
- name: AWS_CA_BUNDLE
value: /etc/ssl/vault-ca/ca.crt
volumeMounts:
- name: cache
mountPath: /var/cache/gocache
- name: vault-ca
mountPath: /etc/ssl/vault-ca
readOnly: true
securityContext:
runAsUser: 65532
runAsGroup: 65532
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: "2"
memory: 2Gi
- name: nginx
image: docker.io/nginx:1.29.8-alpine
imagePullPolicy: IfNotPresent
# Bypass the image entrypoint: its config scripts write to a read-only rootfs.
command:
- nginx
- -g
- daemon off;
ports:
- containerPort: 9090
name: gocache
protocol: TCP
volumeMounts:
- name: nginx-config
mountPath: /etc/nginx/nginx.conf
subPath: nginx.conf
readOnly: true
- name: tmp
mountPath: /tmp
securityContext:
runAsUser: 101
runAsGroup: 101
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 500m
memory: 128Mi
volumes:
# Staging cache in front of S3: losing it costs a repopulate, not data.
- name: cache
emptyDir:
sizeLimit: 20Gi
- name: nginx-config
configMap:
name: gocache-nginx
- name: tmp
emptyDir: {}
- name: vault-ca
secret:
secretName: vault-ca-cert
items:
- key: ca.crt
path: ca.crt
+33
View File
@@ -0,0 +1,33 @@
---
# Shared Go build cache (GOCACHEPROG) for CI and developer laptops. Lives in the
# woodpecker namespace because CI is the primary consumer and reads the Secret here.
apiVersion: ceph.unkin.net/v1alpha1
kind: ObjectStoreUser
metadata:
name: gocache
namespace: woodpecker
spec:
displayName: "Go build cache owner"
uid: gocache
maxBuckets: 1
secretName: gocache-s3
retainOnDelete: false
---
apiVersion: ceph.unkin.net/v1alpha1
kind: Bucket
metadata:
name: gocache
namespace: woodpecker
spec:
bucketName: gocache
ownerRef: gocache
versioning: false
# No placementTarget: default (replicated) placement, not the ec target the
# backup buckets use — a build cache is millions of small objects.
tags:
app: gocache
purpose: go-build-cache
retainOnDelete: false
# A cache bucket is never empty, and the operator refuses to delete a
# non-empty bucket without this, wedging the finalizer.
purgeOnDelete: true
+4
View File
@@ -7,6 +7,10 @@ resources:
- cnpg_cluster.yaml
- cnpg_backup.yaml
- cnpg_pooler.yaml
- gocache_bucket.yaml
- configmap_gocache-nginx.yaml
- deployment_gocache.yaml
- service_gocache.yaml
- serviceaccount_arrproxy_ci.yaml
- serviceaccount_autobackup_operator_ci.yaml
- serviceaccount_ghp.yaml
+23
View File
@@ -0,0 +1,23 @@
---
apiVersion: v1
kind: Service
metadata:
name: gocache
namespace: woodpecker
annotations:
purelb.io/addresses: 198.18.200.11
purelb.io/service-group: common
labels:
app.kubernetes.io/name: gocache
spec:
type: LoadBalancer
# Cluster SNATs off-node traffic to a node address, which would defeat the
# nginx allow rules; Local preserves the wireguard client IP.
externalTrafficPolicy: Local
selector:
app.kubernetes.io/name: gocache
ports:
- name: gocache
port: 9090
targetPort: gocache
protocol: TCP
@@ -0,0 +1,6 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ../../../base/haproxy
@@ -10,7 +10,7 @@ resources:
helmCharts:
- name: victoria-metrics-operator
repo: https://victoriametrics.github.io/helm-charts/
version: "0.57.1"
version: "0.67.3"
releaseName: victoria-metrics-operator
namespace: vm-system
valuesFile: values.yaml
+1
View File
@@ -29,6 +29,7 @@ spec:
- path: apps/overlays/*/ghp
- path: apps/overlays/*/gitea
- path: apps/overlays/*/grafana-system
- path: apps/overlays/*/haproxy
- path: apps/overlays/*/inteldeviceplugins-system
- path: apps/overlays/*/jfrog
- path: apps/overlays/*/k8up-system
+2
View File
@@ -43,6 +43,8 @@ spec:
server: https://kubernetes.default.svc
- namespace: 'gitea'
server: https://kubernetes.default.svc
- namespace: 'haproxy'
server: https://kubernetes.default.svc
- namespace: 'jfrog'
server: https://kubernetes.default.svc
- namespace: 'kanidm'