c39af2f9c3
Rework the logging pipeline around a durable message bus so logs survive a ClickHouse outage, can be replayed after a bad transform, and fan out to multiple independent consumers. Add long-term raw-log backup to S3. Topology becomes edge -> JetStream -> consumers -> sinks: - Dedicated JetStream NATS cluster (3 replicas, file storage) in the logging namespace. Deliberately separate from app messaging (streamstack) for blast-radius isolation. Stream LOGS (subjects logs.>, retention=limits, 40GiB / 72h) is the outage buffer; durable consumers give independent offsets. - Edge publishers (thin): the k8s DaemonSet and a new VM-ingest Deployment (HTTP NDJSON front door behind the logs-ingest Gateway) publish into JetStream (logs.k8s.<ns>.<container> / logs.vm.<host>). No parsing on the edge. - Transform tier (StatefulSet): pulls the whole stream via the durable `transform` consumer, routes by subject, shapes, and remains the sole ClickHouse writer. Its disk buffer shrinks (JetStream is the outage buffer). - Archiver (Deployment): its OWN durable `archiver` consumer (independent offsets — archive lag never affects the ClickHouse path) writes RAW, pre-transform events to a Ceph RGW S3 bucket (cephrgw-operator ObjectStoreUser + Bucket + BucketAccess) as gzipped NDJSON keyed by raw/<subject>/YYYY/MM/DD/. Default subject filter is Vault audit (logs.k8s.vault.>), configurable. Auth: distinct NATS users (producer publish-only, consumer pull+ack, admin for the stream/consumer bootstrap Job) with passwords from Vault (nats-auth Secret); S3 creds from the BucketAccess Secret. Streams/consumers are provisioned by an idempotent PostSync bootstrap Job. Add local kubeconform schemas for the ceph.unkin.net CRDs (datreeio lacks them) and extend the vector-test CI to cover the agent, VM-ingest and archiver configs. Verified end-to-end locally: NATS ACLs, vector JetStream publish, and durable-consumer pull+ack (at-least-once) all work. Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
63 lines
1.8 KiB
YAML
63 lines
1.8 KiB
YAML
---
|
|
# Vector ARCHIVER tier — long-term raw-log backup to S3 (Ceph RGW). Independent
|
|
# durable JetStream consumer (`archiver`) so its offsets/lag are fully isolated
|
|
# from the ClickHouse transform path (archive lag can never stall ingest — true
|
|
# fan-out). Writes RAW, pre-transform events (as they sit in JetStream) as
|
|
# gzipped NDJSON, partitioned by subject + date. This is the long-horizon replay
|
|
# source beyond JetStream's 72h retention window.
|
|
data_dir: /vector-data-dir
|
|
|
|
api:
|
|
enabled: true
|
|
address: 0.0.0.0:8686
|
|
|
|
sources:
|
|
js_archive:
|
|
type: nats
|
|
url: nats://nats.logging.svc.cluster.local:4222
|
|
connection_name: vector-archiver
|
|
subject: "logs.>"
|
|
jetstream:
|
|
stream: LOGS
|
|
consumer: archiver
|
|
auth:
|
|
strategy: user_password
|
|
user_password:
|
|
user: log-consumer
|
|
password: ${NATS_CONSUMER_PASSWORD}
|
|
decoding:
|
|
codec: json
|
|
|
|
sinks:
|
|
s3:
|
|
type: aws_s3
|
|
inputs:
|
|
- js_archive
|
|
bucket: logs-archive
|
|
endpoint: https://s3.ceph.unkin.net
|
|
region: us-east-1
|
|
force_path_style: true
|
|
tls:
|
|
ca_file: /etc/vault-ca/ca.crt
|
|
# AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY come from the logs-archive-s3
|
|
# Secret (cephrgw-operator) via envFrom on the deployment.
|
|
key_prefix: "raw/{{ subject }}/%Y/%m/%d/"
|
|
compression: gzip
|
|
encoding:
|
|
codec: json
|
|
framing:
|
|
method: newline_delimited
|
|
filename_time_format: "%Y%m%dT%H%M%SZ"
|
|
filename_append_uuid: true
|
|
batch:
|
|
max_bytes: 134217728
|
|
timeout_secs: 300
|
|
buffer:
|
|
type: memory
|
|
max_events: 5000
|
|
when_full: block
|
|
# Disabled so slow BucketAccess credential propagation doesn't crash-loop
|
|
# the pod; RGW reachability is proven by the operator's own health.
|
|
healthcheck:
|
|
enabled: false
|