## Why
After v0.1.4 pointed HA peer URLs at per-pod ClusterIP Services, kea-dhcp4 2.6
starts the HA service but then crashes at hook load:
DHCP4_CONFIG_LOAD_FAIL ... Error initializing hooks: CmdHttpListener::run
failed: unable to setup TCP acceptor for listening to the incoming HTTP
requests: bind: Cannot assign requested address
With core multi-threading enabled (Kea 2.6 default), the HA hook opens a
dedicated HTTP listener bound to *this* server's peer url address. That address
is now a per-pod ClusterIP — virtual (kube-proxy DNAT), not assignable on the
pod — so the bind fails. Peers must be reached via ClusterIP, but the local
listener must bind a pod-local address.
## How
- Set the HA relationship's multi-threading block http-dedicated-listener:false
(enable-multi-threading:true). Per Kea 2.6 docs this makes inbound HA traffic
flow through kea-ctrl-agent instead of a hook-owned listener. The ctrl-agent
sidecar already binds 0.0.0.0:8000, and the per-pod Service targetPort 8000
routes ClusterIP:8000 to that container, so remote peers keep reaching this
server at its ClusterIP while nothing binds the virtual address locally.
- Regression tests: rendered config disables the dedicated listener (string +
parsed high-availability.multi-threading assertion); ctrl-agent binds 0.0.0.0
(the pod-local address the CA-mediated route depends on).
## Why
kea-dhcp4 crash-loops at HA hook load: kea 2.6's HA hook parses each peer url host as an IP literal and never resolves DNS, so the StatefulSet headless hostnames are rejected ("Failed to convert string to address ..."). Verified in-cluster that only an IP works (short name, FQDN both fail; `kea-dhcp4 -t` does not exercise this, which is why the v0.1.3 wait did not catch it). Pod IPs cannot be baked into the config because they change on restart and would roll-loop the StatefulSet via the config hash.
## How
- create one ClusterIP Service per HA peer, selecting the pod by its statefulset.kubernetes.io/pod-name label, with publishNotReadyAddresses so peers are routable during bootstrap
- render each HA peer url as its peer Service ClusterIP (a stable IP literal, safe in the config hash); reconcile Services before the ConfigMap and requeue until the ClusterIPs are allocated
Container entrypoints were rendered as `fmt.Sprintf` shell strings in Go, so
every entrypoint fix needed a full operator release, nothing was shellcheckable,
and the escaping was a hazard.
How:
- Add a `kea-init` initContainer that finalises the per-pod config and
bounded-waits for HA peer DNS, replacing the in-entrypoint retry. It hardens
the shared run dir to 0750, substitutes `this-server-name` from the pod
ordinal (`POD_NAME` via the downward API), stages both configs into the shared
emptyDir, and gates on `kea-dhcp4 -t` (60x2s) — failing loud after the cap so
the kubelet restarts it instead of starting a doomed server.
- Run the main kea-dhcp4 / kea-ctrl-agent containers with kea exec'd directly,
dropping both wrapper shells.
- Replace the two `fmt.Sprintf` entrypoints with a single committed
`internal/kea/scripts/init.sh` embedded via `go:embed` and parameterised
entirely by env vars — no Go string interpolation.
- Add a shellcheck step to the pre-commit pipeline.
Test:
- Assert the pod shape: one kea-init initContainer, POD_NAME from the downward
API, main containers exec kea directly, and the ConfigMap carries init.sh (not
the old per-container entrypoints).
- Assert init.sh hardens the socket dir, gates on `kea-dhcp4 -t`, fails loud
after the cap, and is free of fmt verbs.
- shellcheck the embedded script.
## Why
kea-dhcp4 crash-loops on a cold container start: the HA hook resolves the StatefulSet peer URL hostnames once at config load, but the peer DNS records are not resolvable in the first moment of a fresh container, and kea exits hard instead of retrying (HA_CONFIGURATION_FAILED / "Failed to convert string to address"). Once DNS is warm the exact config validates, so the failure is purely a startup race.
## How
- gate the dhcp4 entrypoint on `kea-dhcp4 -t` and retry until the config validates before exec'ing the server
After v0.1.1 moved the socket dir to /var/run/kea, kea-dhcp4 and
kea-ctrl-agent still crash-loop:
DHCP4_PARSER_COMMIT_FAIL ... 'socket-name' is invalid: socket path:/var/run/kea
does not exist or has more relaxed permissions than 750
Kea 2.6+ refuses a unix-socket directory whose mode is more relaxed than
0750. The shared emptyDir is mounted at /var/run/kea with the default 0777,
so kea rejects it. The kea containers run as root, so the entrypoints can
tighten it.
- chmod 0750 the RunDir in both rendered entrypoints after mkdir.
- Assert both entrypoints chmod the socket dir to 0750.
Needs a v0.1.2 release so argocd-apps can bump the operator image.
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
Kea 2.6.5 restricts control/HA unix socket paths to its compiled
runstatedir and rejects any other path by exact string match
("invalid path specified: '/run/kea', supported path is '/var/run/kea'"),
even though /var/run is a symlink to /run. The operator rendered sockets
under /run/kea, so kea-dhcp4 and kea-ctrl-agent crash-looped on startup.
- Point RunDir at /var/run/kea so all derived config/socket paths match.
- Pre-create /var/run/kea in the kea image.
- Assert rendered socket paths live under /var/run/kea.
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
Two sample bugs surfaced while authoring the argocd deployment.
- set subnet gateways to match the authoritative puppet hieradata: 198.18.13-16
routers are .254 (not .1); 198.18.17 stays .1
- rename KeaClientClass samples Legacy/UEFI-64 to legacy/uefi-64 so they are
valid RFC1123 object names (the operator renders the kea class name from
metadata.name, which k8s forces to lowercase); align the README
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
The 'local' subnet is not needed in the kea migration. Remove it from the ISC
translation sample; the pool-less-subnet capability and its unit tests remain.
- remove net-198-18-25 KeaSubnet from config/samples/01-subnets.yaml
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
Replace the ISC dhcpd PXE-boot VM with a Kea DHCP Kubernetes operator, modelled
on bind-operator. The operator renders kea-dhcp4 config from CRs and runs an HA
pair of kea-dhcp4 + kea-ctrl-agent servers behind an anycast Service.
- add KeaCluster/KeaSubnet/KeaClientClass/KeaAPI CRDs (group kea.unkin.net)
- render deterministic kea-dhcp4.conf + kea-ctrl-agent.conf into a ConfigMap and
roll the StatefulSet via a config-hash annotation; best-effort hot-reload via
the kea-ctrl-agent REST channel
- run HA hot-standby (memfile leases) with stable per-peer DNS identity from a
StatefulSet; expose an anycast LoadBalancer Service for PureLB
- represent the full legacy dhcpd config: 198.18.13-17.0/24 pools, pool-less
198.18.25.0/24, and the Legacy/UEFI-64 PXE arch classes (option 93)
- add the KeaAPI-spawned REST service: Terraform-friendly CRUD over subnet and
client-class CRs (stable IDs, PUT upsert, 404 drift, bearer-token auth)
- add Makefile (patch/minor/major tag targets), distroless operator/api images,
an AlmaLinux+EPEL kea workload image, and woodpecker CI with k8s resources +
serviceAccountName on every step
- unit tests for config rendering, controller reconcile/config-hash, and the API
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT