10020033d9
Stand up a centralized logging estate that captures ALL logs from k8s pods and (via a reachable ingestion endpoint) puppet-managed VMs, storing them in ClickHouse for query/retention. Metrics already live in VictoriaMetrics; this adds the logs pillar under a dedicated `logging` ArgoCD project. Deploy the Altinity clickhouse-operator (clickhouse-system) and a single-shard ClickHouseInstallation (logging) on cephrbd-fast-delete with a MergeTree logs.raw table (30d TTL) bootstrapped by an idempotent PostSync Job. Deploy Vector as an explicit two tiers: - Edge (thin): a DaemonSet tails every node's pod logs and forwards over the Vector-native protocol to the aggregator; no parsing at the edge. Future VM agents follow the same thin pattern. - Aggregator (brain): HA StatefulSet that is the sole ClickHouse writer, holds the only ClickHouse credentials, owns all transforms, batches into few fat inserts (avoid too-many-parts), and buffers to disk (PVC) to ride out a ClickHouse outage. Its pipeline is a single source-of-truth config validated by `vector test` in CI; per-app pipelines become aggregator-only changes. Expose the VM ingestion endpoint at logs-ingest.k8s.syd1.au.unkin.net via the internal Traefik gateway (cert-manager + external-dns), routing to the aggregator's HTTP source so puppet VMs can reach it over TLS. Source ClickHouse credentials from Vault via the existing VaultStaticSecret pattern (templated k8s auth policy already grants the logging namespace); password hash never lands in git. Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv