No agent Vault role covers the VictoriaMetrics/VictoriaLogs stack, so a scoped Kubernetes token cannot be issued for it and writes there fall back to an admin context.
- add the agent-observability kubernetes secret backend role, allowed in vm-system, observability and logging
- add its generated role rules: read plus patch/update on VictoriaMetrics CRs and workloads, pod delete for rolling restarts, read-only on services, configmaps, endpoints, events and Gateway API routes
- add a policy granting update on kubernetes/au/syd1/creds/agent-observability to the cluster_operator LDAP group and the agents approle
Reviewed-on: #156
Co-authored-by: unkin-agent <unkin-agent@unkin.net>
Co-committed-by: unkin-agent <unkin-agent@unkin.net>
## Why
The `operator` kube context is a Vault-minted, read-only credential (Kubernetes
secret engine role `cluster-operator`, bound to a `get/list/watch`-only
ClusterRole). It is currently RBAC-forbidden from listing operator-owned CRDs —
the immediate breakage is `valkeyclusters.valkey.io` — and likewise every other
operator CRD group deployed via `argocd-apps`. This extends the RO ruleset so the
context can read those CRDs. Still strictly read-only: no create/update/delete.
## Change
- Extend the `cluster-operator` generated_role_rules
(`resources/secret_backend/kubernetes/au/syd1/roles/cluster-operator.yaml`)
with `get/list/watch` on the CRD API groups of the operators deployed via
`argocd-apps` (verbs and `resources: "*"` unchanged; same single rule block).
## API groups added
- `valkey.io` (valkey-operator — immediate need)
- `ceph.unkin.net` (cephrgw-operator)
- `bind.unkin.net` (bind-operator)
- `kea.unkin.net` (kea/dhcp operator)
- `k8up.io` (k8up)
- `grafana.integreatly.org` (grafana-operator)
- `operator.victoriametrics.com` (VictoriaMetrics operator)
- `clickhouse.altinity.com`, `clickhouse-keeper.altinity.com` (altinity clickhouse-operator)
- `acme.cert-manager.io` (cert-manager companion CRD group)
- `deviceplugin.intel.com`, `fpga.intel.com` (intel device plugins operator)
- `autoscaling.k8s.io` (VPA)
- `apm.k8s.elastic.co`, `beat.k8s.elastic.co`, `agent.k8s.elastic.co`,
`maps.k8s.elastic.co`, `enterprisesearch.k8s.elastic.co`,
`autoscaling.k8s.elastic.co`, `stackconfigpolicy.k8s.elastic.co` (ECK — the
`elasticsearch`/`kibana`/`logstash` ECK groups were already granted)
- `snapshot.storage.k8s.io`, `groupsnapshot.storage.k8s.io` (CSI external-snapshotter, deployed via csi-cephfs/csi-cephrbd)
Groups already present (`postgresql.cnpg.io`, `cert-manager.io`,
`externaldns.k8s.io`, `secrets.hashicorp.com`, `purelb.io`, `nfd.k8s-sigs.io`,
`elasticsearch/kibana/logstash.k8s.elastic.co`, `gateway.networking.k8s.io`,
etc.) are unchanged. Rancher/RKE/Calico/cluster-api/fleet management-layer CRD
groups are intentionally excluded — they are not `argocd-apps` operators.
---------
Co-authored-by: unkin-agent <agent@unkin.net>
Reviewed-on: #135
Co-authored-by: Unkin Agent <unkin-agent@unkin.net>
Co-committed-by: Unkin Agent <unkin-agent@unkin.net>
## Why
Agentic workloads currently need cluster-admin/root kubeconfig contexts to do routine per-domain work. This adds domain-scoped, Vault-issued kubernetes credentials plus an `agents` AppRole so agents get least-privilege access instead of escalating.
## Changes
- Add `kubernetes_secret_backend_role` configs `agent-dhcp`, `agent-dns`, `agent-certs`, `agent-storage` (au/syd1):
- **agent-dhcp** (Role, ns `dhcp-system`): full verbs on `kea.unkin.net` CRDs; get/list/watch pods/services/configmaps/events + pods/log.
- **agent-dns** (`service_account_name` mode): mints tokens for the static `agent-dns` SA (argocd-apps#332) whose per-namespace RoleBindings confine access to bind-system/bind-internal/bind-external/externaldns. `allowed_kubernetes_namespaces: [bind-system]` (the SA's namespace).
- **agent-certs** (Role, ns `cert-manager`): full verbs on `cert-manager.io` + `acme.cert-manager.io` (closes the orders/challenges debugging gap); get/list/watch/delete secrets; get/list/watch pods + pods/log. Secret delete confined to `cert-manager`.
- **agent-storage** (Role, ns `cephrgw-system`): full verbs on `ceph.unkin.net` CRDs (buckets/bucketaccesses/objectstoreusers); get/list/watch pods + pods/log.
- Extend the `kubernetes_secret_backend_role` module with an optional `service_account_name`; when set, `generated_role_rules`/`kubernetes_role_type` are omitted (the SA's own bindings supply RBAC).
- Add creds policies for each role, bound to the `kubernetes_au_syd1_cluster_operator` ldap group (human kubectl use) and the `agents` AppRole (programmatic use).
- Add the `agents` AppRole (mirrors the certmanager approle schema): `bind_secret_id: false` (role_id-only login), `token_bound_cidrs: [10.10.12.200/32]` (agent workstation wg0 addr), deterministic role_id, 1h/4h TTLs.
- Add `policies/kv/kubernetes/agents.yaml` granting the AppRole create/read/update/list on `kv/data/kubernetes/*` + read/list on `kv/metadata/kubernetes/*` (no delete).
## Ordering
argocd-apps#332 (the `agent-dns` SA + ClusterRole + per-namespace RoleBindings) must sync **before** the `agent-dns` creds here are usable — Vault mints tokens for a service account that must already exist.
https://claude.ai/code/session_01JUoARVdmhxKQHyyyp1pxeT
Reviewed-on: #109
Co-authored-by: Ben Vincent <ben@unkin.net>
Co-committed-by: Ben Vincent <ben@unkin.net>
- migrate from individual terraform files to config-driven terragrunt module structure
- add vault_cluster module with config discovery system
- replace individual .tf files with centralized config.hcl
- restructure auth and secret backends as configurable modules
- move auth roles and secret backends to yaml-based configuration
- convert policies from .hcl to .yaml format, add rules/auth definition
- add pre-commit hooks for yaml formatting and file cleanup
- add terragrunt cache to gitignore
- update makefile with terragrunt commands and format target