## Why
k8s nodes fail every puppet run on Package[rke2-server] and stop
applying the rest of their catalog (no further package/config updates):
change from '1.33.4~rke2r1-1.el9' to '1.33.11~rke2r1' failed:
dnf upgrade rke2-server-1.33.11~rke2r1 returned 1:
Problem: problem with installed package rke2-common-1.33.13~rke2r2-0.el9
- package rke2-server-1.33.11~rke2r1 requires rke2-common = 1.33.11~rke2r1,
but none of the providers can be installed
- cannot install the best update candidate for rke2-server
The module versionlocks only rke2-server/rke2-agent, not their strict
(= version) rke2-common dependency. rke2-common is pulled from the
rolling rancher-rke2-1.33-latest channel, whose head is now
1.33.13~rke2r2, so rke2-common drifted up to 1.33.13~rke2r2 while the
pin sat at 1.33.11~rke2r1. dnf upgrade can't downgrade the newer
rke2-common to satisfy the older server, so the transaction fails.
## Changes
- Versionlock rke2-common to the same ${rke2_version}~${rke2_release} as
the server/agent, so the rolling channel can no longer drift the
dependency ahead of the pin.
- Bump rke2_version 1.33.11 -> 1.33.13 and rke2_release rke2r1 -> rke2r2
to match the current channel head (and the already-drifted installed
rke2-common), so the pinned server/agent, versionlocks, and preloaded
airgap bundle all resolve in one transaction.
## Why
Four newly-provisioned el9_8 compute nodes (prodnxsr0014/0015/0018/0019) hang with canal's kube-flannel container in `ImagePullBackOff`: the flannel VXLAN overlay never comes up, so the node can't reach any in-cluster `198.18.200.x` LoadBalancer VIP.
Root cause is a stale airgap-preload version. The nodes boot rke2 **v1.33.11+rke2r1** from the rolling `latest/1.33` repo, whose canal manifest requests `hardened-flannel:v0.28.4` / `hardened-calico:v3.31.5` (build20260415). But `rke2::install` pinned the preloaded bundle to **v1.33.4+rke2r1** (flannel v0.27.2 / calico v3.30.2), so those images were never on disk. containerd then falls back to the `docker.io` mirror (artifactapi, `disable-default-registry-endpoint: true`), reachable only via the pod-overlay VIP that requires the very flannel image being pulled — a bootstrap deadlock. Older nodes (0001-0008) are versionlocked at 1.33.4 and still match their original bundle, so they were unaffected.
## Changes
- Bump pinned `rke2_version` `1.33.4` -> `1.33.11` so the versionlock, RPM `ensure`, and preloaded bundle all line up with the canal image tags the running binary requests. The default `rke2-images.linux-amd64.tar.zst` bundle already contains the canal CNI images (it is RKE2's default CNI), so no extra tarball is needed.
- Wire the airgap archive `source` to the `container_archive_source` class parameter (previously declared in the module but never consumed). The module keeps its generic upstream default; the artifactapi override (the pre-CNI-reachable source, same BGP/physical path the rke2 yum repos already use) lives in the k8s role hiera as `rke2::container_archive_source`.
Applies to servers and agents alike (`rke2::install` runs for both) and preloads `before => Service`, so the bundle lands before rke2 starts.
Notes:
- The `latest/1.33` repo is rolling, so the pinned version must be maintained as the repo advances; a follow-up to pin the yum channel to a fixed patch would remove the drift entirely.
- No terraform-artifactapi change is required. (If a canal-only preload were ever wanted, the github generic remote allowlist would need `rancher/rke2/.*/rke2-images-canal.linux-amd64.tar.zst$` added — but the default bundle already carries those images, so it is unnecessary.)
https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
Reviewed-on: #512
Co-authored-by: Ben Vincent <ben@unkin.net>
Co-committed-by: Ben Vincent <ben@unkin.net>