ceph: manage /etc/ceph/ceph.conf on osd and mon/mgr/mds hosts #515

Merged
benvin merged 3 commits from benvin/manage-ceph-conf into develop 2026-08-09 00:00:21 +10:00

3 Commits

Author SHA1 Message Date
unkinben fa3b97058c ceph: include prodnxsr0014-0019 in public_network
ci/woodpecker/pr/ruby-validate Pipeline was successful
ci/woodpecker/pr/puppet-lint Pipeline was successful
ci/woodpecker/pr/bolt-validate Pipeline was successful
ci/woodpecker/pr/yamllint Pipeline was successful
ci/woodpecker/pr/erb-validate Pipeline was successful
ci/woodpecker/pr/epp-validate Pipeline was successful
ci/woodpecker/pr/puppet-validate Pipeline was successful
ci/woodpecker/pr/ruby-check Pipeline was successful
Why: prodnxsr0014-0019 were racked as roles::base but have since joined
the de96a98f ceph cluster as OSD hosts (live enc_role is now
roles::infra::k8s::compute, is_ceph_osd true). Their ceph-public /32s
must appear in public_network on every cluster member, and they must
receive the osd-only ceph.conf.

How:
- Expand profiles::ceph::client::cluster_public_ips from .1-.13 to
  .1-.19 (add ceph-public loopbacks 198.18.23.14-19).
- No role/hiera change needed for 0014-0019: they run
  roles::infra::k8s::compute, already covered by roles/infra/k8s.yaml
  (manage_ceph_conf true, render_mds_config unset/false), so they get
  the [global]-only variant with no mds sections.

Consequence: first convergence now rewrites public_network on the
existing osd hosts (0001-0008) and mon hosts (0009-0013) too, adding
.14-.19; adoption is a single public_network-line change on every
cluster host rather than a no-op on osd hosts. Verified against the
live files: osd (0008) and mon (0009) each differ by exactly the
public_network line, mds formatting on 0009 unchanged.
2026-08-08 22:25:53 +10:00
unkinben 5b04aa341d ceph: match live blank-line separators between mds sections
ci/woodpecker/pr/ruby-validate Pipeline was successful
ci/woodpecker/pr/puppet-lint Pipeline was successful
ci/woodpecker/pr/erb-validate Pipeline was successful
ci/woodpecker/pr/bolt-validate Pipeline was successful
ci/woodpecker/pr/yamllint Pipeline was successful
ci/woodpecker/pr/epp-validate Pipeline was successful
ci/woodpecker/pr/ruby-check Pipeline was successful
ci/woodpecker/pr/puppet-validate Pipeline was successful
Diffing the render against the live /etc/ceph/ceph.conf on prodnxsr0009
showed the hand-maintained file separates the mds sections with blank
lines: one before [mds] and one before each [mds.X-i] block (no trailing
blank line after the last). Emit those separators so a mon/mgr/mds host
renders byte-identical to live except for the intended public_network
normalization. Verified: osd host stays byte-identical, mon host differs
only on the public_network line.

Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
2026-08-08 20:04:04 +10:00
unkinben 2d125487c3 ceph: manage /etc/ceph/ceph.conf on osd and mon/mgr/mds hosts
ci/woodpecker/pr/ruby-validate Pipeline was successful
ci/woodpecker/pr/erb-validate Pipeline was successful
ci/woodpecker/pr/puppet-lint Pipeline was successful
ci/woodpecker/pr/bolt-validate Pipeline was successful
ci/woodpecker/pr/yamllint Pipeline was successful
ci/woodpecker/pr/epp-validate Pipeline was successful
ci/woodpecker/pr/ruby-check Pipeline was successful
ci/woodpecker/pr/puppet-validate Pipeline was successful
Bring the hand-maintained /etc/ceph/ceph.conf under Puppet on the
prodnxsr ceph cluster (fsid de96a98f). The file is identified per host
by role: k8s roles include profiles::ceph::osd only and get a [global]
section; the incus node role also includes profiles::ceph::mds and
additionally gets the [mds] + [mds.*] sections.

- Add cluster topology as a single source of truth in common.yaml:
  cluster_public_ips (all 13 ceph host /32s), mon_initial_members, and
  the mds_instances map (two mds daemons per mon/mgr/mds host).
- Render /etc/ceph/ceph.conf from that topology in the reworked
  client.conf.erb, preserving the live two-space indent and key order.
- Gate the [mds] sections on render_mds_config, set true only in the
  incus node role hiera (the role that includes profiles::ceph::mds).
- Drop the hard Package[ceph-common] dependency when the class does not
  manage the package (cephadm/profiles::packages deliver it on the k8s
  and incus hosts).
- Enable manage_ceph_conf on the k8s and incus node roles.

public_network is normalized to all 13 ceph host /32s on every host;
this rewrites it on the mon/mgr/mds hosts (adding .1-.8) as the one
intended content change. RGW hosts keep their profiles::ceph::conf
variant and are untouched.

Claude-Session: https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
2026-08-08 19:56:49 +10:00