6 Commits

Author SHA1 Message Date
unkin-agent b0e3e31a6b Disable HA hook dedicated listener; route HA via ctrl-agent
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/test Pipeline was successful
ci/woodpecker/pr/build Pipeline was successful
## Why
After v0.1.4 pointed HA peer URLs at per-pod ClusterIP Services, kea-dhcp4 2.6
starts the HA service but then crashes at hook load:

  DHCP4_CONFIG_LOAD_FAIL ... Error initializing hooks: CmdHttpListener::run
  failed: unable to setup TCP acceptor for listening to the incoming HTTP
  requests: bind: Cannot assign requested address

With core multi-threading enabled (Kea 2.6 default), the HA hook opens a
dedicated HTTP listener bound to *this* server's peer url address. That address
is now a per-pod ClusterIP — virtual (kube-proxy DNAT), not assignable on the
pod — so the bind fails. Peers must be reached via ClusterIP, but the local
listener must bind a pod-local address.

## How
- Set the HA relationship's multi-threading block http-dedicated-listener:false
  (enable-multi-threading:true). Per Kea 2.6 docs this makes inbound HA traffic
  flow through kea-ctrl-agent instead of a hook-owned listener. The ctrl-agent
  sidecar already binds 0.0.0.0:8000, and the per-pod Service targetPort 8000
  routes ClusterIP:8000 to that container, so remote peers keep reaching this
  server at its ClusterIP while nothing binds the virtual address locally.
- Regression tests: rendered config disables the dedicated listener (string +
  parsed high-availability.multi-threading assertion); ctrl-agent binds 0.0.0.0
  (the pod-local address the CA-mediated route depends on).
2026-08-27 00:21:20 +10:00
unkinben 9b3fa83dac Move kea entrypoints out of Go fmt.Sprintf into an initContainer
ci/woodpecker/pr/build Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/test Pipeline was successful
Container entrypoints were rendered as `fmt.Sprintf` shell strings in Go, so
every entrypoint fix needed a full operator release, nothing was shellcheckable,
and the escaping was a hazard.

How:
- Add a `kea-init` initContainer that finalises the per-pod config and
  bounded-waits for HA peer DNS, replacing the in-entrypoint retry. It hardens
  the shared run dir to 0750, substitutes `this-server-name` from the pod
  ordinal (`POD_NAME` via the downward API), stages both configs into the shared
  emptyDir, and gates on `kea-dhcp4 -t` (60x2s) — failing loud after the cap so
  the kubelet restarts it instead of starting a doomed server.
- Run the main kea-dhcp4 / kea-ctrl-agent containers with kea exec'd directly,
  dropping both wrapper shells.
- Replace the two `fmt.Sprintf` entrypoints with a single committed
  `internal/kea/scripts/init.sh` embedded via `go:embed` and parameterised
  entirely by env vars — no Go string interpolation.
- Add a shellcheck step to the pre-commit pipeline.

Test:
- Assert the pod shape: one kea-init initContainer, POD_NAME from the downward
  API, main containers exec kea directly, and the ConfigMap carries init.sh (not
  the old per-container entrypoints).
- Assert init.sh hardens the socket dir, gates on `kea-dhcp4 -t`, fails loud
  after the cap, and is free of fmt verbs.
- shellcheck the embedded script.
2026-08-08 22:52:59 +10:00
unkinben f20b26fd71 Wait for HA peer DNS before starting kea-dhcp4
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/test Pipeline was successful
ci/woodpecker/pr/build Pipeline was successful
## Why
kea-dhcp4 crash-loops on a cold container start: the HA hook resolves the StatefulSet peer URL hostnames once at config load, but the peer DNS records are not resolvable in the first moment of a fresh container, and kea exits hard instead of retrying (HA_CONFIGURATION_FAILED / "Failed to convert string to address"). Once DNS is warm the exact config validates, so the failure is purely a startup race.

## How
- gate the dhcp4 entrypoint on `kea-dhcp4 -t` and retry until the config validates before exec'ing the server
2026-08-08 22:38:20 +10:00
unkinben e95e5437a2 Harden kea socket dir to 0750 in the rendered entrypoints
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/test Pipeline was successful
ci/woodpecker/pr/build Pipeline was successful
After v0.1.1 moved the socket dir to /var/run/kea, kea-dhcp4 and
kea-ctrl-agent still crash-loop:

  DHCP4_PARSER_COMMIT_FAIL ... 'socket-name' is invalid: socket path:/var/run/kea
  does not exist or has more relaxed permissions than 750

Kea 2.6+ refuses a unix-socket directory whose mode is more relaxed than
0750. The shared emptyDir is mounted at /var/run/kea with the default 0777,
so kea rejects it. The kea containers run as root, so the entrypoints can
tighten it.

- chmod 0750 the RunDir in both rendered entrypoints after mkdir.
- Assert both entrypoints chmod the socket dir to 0750.

Needs a v0.1.2 release so argocd-apps can bump the operator image.

Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
2026-08-08 20:12:17 +10:00
unkinben 1b8be63d05 Render kea unix socket paths under /var/run/kea
ci/woodpecker/pr/pre-commit Pipeline was successful
ci/woodpecker/pr/test Pipeline was successful
ci/woodpecker/pr/build Pipeline was successful
Kea 2.6.5 restricts control/HA unix socket paths to its compiled
runstatedir and rejects any other path by exact string match
("invalid path specified: '/run/kea', supported path is '/var/run/kea'"),
even though /var/run is a symlink to /run. The operator rendered sockets
under /run/kea, so kea-dhcp4 and kea-ctrl-agent crash-looped on startup.

- Point RunDir at /var/run/kea so all derived config/socket paths match.
- Pre-create /var/run/kea in the kea image.
- Assert rendered socket paths live under /var/run/kea.

Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
2026-08-06 23:38:55 +10:00
unkinben d3fb5dcd1a Scaffold kea-operator: CRDs, controllers, config rendering, REST API, CI
ci/woodpecker/pr/build Pipeline failed
ci/woodpecker/pr/pre-commit Pipeline failed
ci/woodpecker/pr/test Pipeline failed
Replace the ISC dhcpd PXE-boot VM with a Kea DHCP Kubernetes operator, modelled
on bind-operator. The operator renders kea-dhcp4 config from CRs and runs an HA
pair of kea-dhcp4 + kea-ctrl-agent servers behind an anycast Service.

- add KeaCluster/KeaSubnet/KeaClientClass/KeaAPI CRDs (group kea.unkin.net)
- render deterministic kea-dhcp4.conf + kea-ctrl-agent.conf into a ConfigMap and
  roll the StatefulSet via a config-hash annotation; best-effort hot-reload via
  the kea-ctrl-agent REST channel
- run HA hot-standby (memfile leases) with stable per-peer DNS identity from a
  StatefulSet; expose an anycast LoadBalancer Service for PureLB
- represent the full legacy dhcpd config: 198.18.13-17.0/24 pools, pool-less
  198.18.25.0/24, and the Legacy/UEFI-64 PXE arch classes (option 93)
- add the KeaAPI-spawned REST service: Terraform-friendly CRUD over subnet and
  client-class CRs (stable IDs, PUT upsert, 404 drift, bearer-token auth)
- add Makefile (patch/minor/major tag targets), distroless operator/api images,
  an AlmaLinux+EPEL kea workload image, and woodpecker CI with k8s resources +
  serviceAccountName on every step
- unit tests for config rendering, controller reconcile/config-hash, and the API

Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
2026-08-02 18:42:46 +10:00