Container entrypoints were rendered as `fmt.Sprintf` shell strings in Go, so
every entrypoint fix needed a full operator release, nothing was shellcheckable,
and the escaping was a hazard.
How:
- Add a `kea-init` initContainer that finalises the per-pod config and
bounded-waits for HA peer DNS, replacing the in-entrypoint retry. It hardens
the shared run dir to 0750, substitutes `this-server-name` from the pod
ordinal (`POD_NAME` via the downward API), stages both configs into the shared
emptyDir, and gates on `kea-dhcp4 -t` (60x2s) — failing loud after the cap so
the kubelet restarts it instead of starting a doomed server.
- Run the main kea-dhcp4 / kea-ctrl-agent containers with kea exec'd directly,
dropping both wrapper shells.
- Replace the two `fmt.Sprintf` entrypoints with a single committed
`internal/kea/scripts/init.sh` embedded via `go:embed` and parameterised
entirely by env vars — no Go string interpolation.
- Add a shellcheck step to the pre-commit pipeline.
Test:
- Assert the pod shape: one kea-init initContainer, POD_NAME from the downward
API, main containers exec kea directly, and the ConfigMap carries init.sh (not
the old per-container entrypoints).
- Assert init.sh hardens the socket dir, gates on `kea-dhcp4 -t`, fails loud
after the cap, and is free of fmt verbs.
- shellcheck the embedded script.
## Why
kea-dhcp4 crash-loops on a cold container start: the HA hook resolves the StatefulSet peer URL hostnames once at config load, but the peer DNS records are not resolvable in the first moment of a fresh container, and kea exits hard instead of retrying (HA_CONFIGURATION_FAILED / "Failed to convert string to address"). Once DNS is warm the exact config validates, so the failure is purely a startup race.
## How
- gate the dhcp4 entrypoint on `kea-dhcp4 -t` and retry until the config validates before exec'ing the server
After v0.1.1 moved the socket dir to /var/run/kea, kea-dhcp4 and
kea-ctrl-agent still crash-loop:
DHCP4_PARSER_COMMIT_FAIL ... 'socket-name' is invalid: socket path:/var/run/kea
does not exist or has more relaxed permissions than 750
Kea 2.6+ refuses a unix-socket directory whose mode is more relaxed than
0750. The shared emptyDir is mounted at /var/run/kea with the default 0777,
so kea rejects it. The kea containers run as root, so the entrypoints can
tighten it.
- chmod 0750 the RunDir in both rendered entrypoints after mkdir.
- Assert both entrypoints chmod the socket dir to 0750.
Needs a v0.1.2 release so argocd-apps can bump the operator image.
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
Replace the ISC dhcpd PXE-boot VM with a Kea DHCP Kubernetes operator, modelled
on bind-operator. The operator renders kea-dhcp4 config from CRs and runs an HA
pair of kea-dhcp4 + kea-ctrl-agent servers behind an anycast Service.
- add KeaCluster/KeaSubnet/KeaClientClass/KeaAPI CRDs (group kea.unkin.net)
- render deterministic kea-dhcp4.conf + kea-ctrl-agent.conf into a ConfigMap and
roll the StatefulSet via a config-hash annotation; best-effort hot-reload via
the kea-ctrl-agent REST channel
- run HA hot-standby (memfile leases) with stable per-peer DNS identity from a
StatefulSet; expose an anycast LoadBalancer Service for PureLB
- represent the full legacy dhcpd config: 198.18.13-17.0/24 pools, pool-less
198.18.25.0/24, and the Legacy/UEFI-64 PXE arch classes (option 93)
- add the KeaAPI-spawned REST service: Terraform-friendly CRUD over subnet and
client-class CRs (stable IDs, PUT upsert, 404 drift, bearer-token auth)
- add Makefile (patch/minor/major tag targets), distroless operator/api images,
an AlmaLinux+EPEL kea workload image, and woodpecker CI with k8s resources +
serviceAccountName on every step
- unit tests for config rendering, controller reconcile/config-hash, and the API
Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT