c05ccfcb5d
logarchiver replaces the plain Vector archiver leg of the centralized logging stack (argocd-apps #296) with a Go service that archives raw logs from NATS JetStream to S3 as zstd-compressed, OpenPGP-encrypted, indexed objects, plus an operator CLI to search the index and retrieve/decrypt archived logs. It adds the things that outgrew Vector: zstd compression, encryption keyed from Ben's Vault GPG secrets engine, a searchable ClickHouse index, and sink-conditional acks (a batch is acknowledged to JetStream only after the object is durably in S3 AND indexed). Service (`logarchiver run`): - Durable JetStream pull consumer (stream LOGS, durable archiver, subject filter default logs.k8s.vault.>), explicit acks, independent offsets. - Batch per subject by size/count/time -> NDJSON -> zstd -> encrypt -> S3 PUT -> ClickHouse index row -> ack. On any failure the batch is Nak'd and redelivered, so nothing is lost on a sink outage. - Encryption is a wrapped-DEK envelope (container LARC1): the bulk is AES-256-GCM framed under a random data key, and only that 32-byte key is OpenPGP-encrypted to the engine's public key. This is because the Vault GPG engine does whole-payload decrypt only; retrieval round-trips just the tiny wrapped key regardless of object size. Public key fetched from the engine or a mounted file (configurable); key fingerprint recorded per object; periodic pubkey refresh for rotation. - Prometheus metrics, structured slog, graceful drain on shutdown. CLI: - `search` queries the index (subject/host/time) and lists matching objects. - `fetch` downloads, decrypts via the Vault GPG engine, unzstds and emits NDJSON (optionally re-filtered by host/time). - `init-schema` creates/prints the ClickHouse archive_index DDL. - cobra `completion` subcommands. Config via file+env (k8s-friendly, secrets from env), boundaries (NATS/S3/ ClickHouse/Vault) behind interfaces with unit tests (config, batching, host/subject extraction, crypto roundtrip with a test key, ack-after-persist with fakes, search query building). go build/vet/test -race clean; golangci-lint v2 clean. Woodpecker CI: build/test/pre-commit on PR; on v* tag a container image plus a Gitea binary release + rpm-internal RPM. Docs per subcommand + architecture + retrieval runbook + deployment drop-in. Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
43 lines
1.6 KiB
Go
43 lines
1.6 KiB
Go
package index
|
|
|
|
import "fmt"
|
|
|
|
// CreateDatabaseSQL creates the index database if absent.
|
|
func CreateDatabaseSQL(database string) string {
|
|
return fmt.Sprintf("CREATE DATABASE IF NOT EXISTS %s", database)
|
|
}
|
|
|
|
// CreateTableSQL returns the DDL for the archive index table. One row is written
|
|
// per archived S3 object. In-cluster the argocd bootstrap Job owns table
|
|
// creation (like the logging stack's clickhouse-schema PostSync hook); this DDL
|
|
// is also shipped as schema/archive_index.sql and applied by `logarchiver
|
|
// init-schema`.
|
|
//
|
|
// PARTITION BY month of min_ts keeps partitions coarse (few objects/day).
|
|
// ORDER BY (subject, min_ts) matches the primary search axes. A bloom_filter
|
|
// skip index on hosts accelerates host lookups without a per-host column.
|
|
func CreateTableSQL(database, table string) string {
|
|
return fmt.Sprintf(`CREATE TABLE IF NOT EXISTS %s.%s
|
|
(
|
|
object_key String,
|
|
bucket LowCardinality(String),
|
|
subject LowCardinality(String),
|
|
hosts Array(LowCardinality(String)),
|
|
min_ts DateTime64(3),
|
|
max_ts DateTime64(3),
|
|
event_count UInt64,
|
|
raw_bytes UInt64,
|
|
stored_bytes UInt64,
|
|
compression LowCardinality(String),
|
|
cipher LowCardinality(String),
|
|
container_format LowCardinality(String),
|
|
key_name LowCardinality(String),
|
|
key_fingerprint String,
|
|
created_at DateTime64(3) DEFAULT now64(3),
|
|
INDEX idx_hosts hosts TYPE bloom_filter GRANULARITY 1
|
|
)
|
|
ENGINE = MergeTree
|
|
PARTITION BY toYYYYMM(min_ts)
|
|
ORDER BY (subject, min_ts, object_key)`, database, table)
|
|
}
|