c05ccfcb5d
logarchiver replaces the plain Vector archiver leg of the centralized logging stack (argocd-apps #296) with a Go service that archives raw logs from NATS JetStream to S3 as zstd-compressed, OpenPGP-encrypted, indexed objects, plus an operator CLI to search the index and retrieve/decrypt archived logs. It adds the things that outgrew Vector: zstd compression, encryption keyed from Ben's Vault GPG secrets engine, a searchable ClickHouse index, and sink-conditional acks (a batch is acknowledged to JetStream only after the object is durably in S3 AND indexed). Service (`logarchiver run`): - Durable JetStream pull consumer (stream LOGS, durable archiver, subject filter default logs.k8s.vault.>), explicit acks, independent offsets. - Batch per subject by size/count/time -> NDJSON -> zstd -> encrypt -> S3 PUT -> ClickHouse index row -> ack. On any failure the batch is Nak'd and redelivered, so nothing is lost on a sink outage. - Encryption is a wrapped-DEK envelope (container LARC1): the bulk is AES-256-GCM framed under a random data key, and only that 32-byte key is OpenPGP-encrypted to the engine's public key. This is because the Vault GPG engine does whole-payload decrypt only; retrieval round-trips just the tiny wrapped key regardless of object size. Public key fetched from the engine or a mounted file (configurable); key fingerprint recorded per object; periodic pubkey refresh for rotation. - Prometheus metrics, structured slog, graceful drain on shutdown. CLI: - `search` queries the index (subject/host/time) and lists matching objects. - `fetch` downloads, decrypts via the Vault GPG engine, unzstds and emits NDJSON (optionally re-filtered by host/time). - `init-schema` creates/prints the ClickHouse archive_index DDL. - cobra `completion` subcommands. Config via file+env (k8s-friendly, secrets from env), boundaries (NATS/S3/ ClickHouse/Vault) behind interfaces with unit tests (config, batching, host/subject extraction, crypto roundtrip with a test key, ack-after-persist with fakes, search query building). go build/vet/test -race clean; golangci-lint v2 clean. Woodpecker CI: build/test/pre-commit on PR; on v* tag a container image plus a Gitea binary release + rpm-internal RPM. Docs per subcommand + architecture + retrieval runbook + deployment drop-in. Claude-Session: https://claude.ai/code/session_015ur3i7D2azsMAWTSVABApv
46 lines
1.4 KiB
Go
46 lines
1.4 KiB
Go
package cli
|
|
|
|
import (
|
|
"fmt"
|
|
|
|
"git.unkin.net/unkin/logarchiver/internal/index"
|
|
"github.com/spf13/cobra"
|
|
)
|
|
|
|
func newInitSchemaCmd() *cobra.Command {
|
|
var printOnly bool
|
|
cmd := &cobra.Command{
|
|
Use: "init-schema",
|
|
Short: "Create the ClickHouse archive-index database and table (idempotent)",
|
|
Long: `init-schema creates the ClickHouse database and archive_index table.
|
|
|
|
In-cluster the argocd bootstrap Job owns schema creation (like the logging
|
|
stack's clickhouse-schema PostSync hook); this command is for local/dev use and
|
|
for emitting the DDL (--print) to embed in that Job.`,
|
|
RunE: func(cmd *cobra.Command, _ []string) error {
|
|
cfg, err := loadConfig()
|
|
if err != nil {
|
|
return err
|
|
}
|
|
if printOnly {
|
|
out := cmd.OutOrStdout()
|
|
_, _ = fmt.Fprintln(out, index.CreateDatabaseSQL(cfg.Index.Database)+";")
|
|
_, _ = fmt.Fprintln(out, index.CreateTableSQL(cfg.Index.Database, cfg.Index.Table)+";")
|
|
return nil
|
|
}
|
|
ch, err := index.NewClickHouse(cmd.Context(), indexConfig(cfg.Index))
|
|
if err != nil {
|
|
return err
|
|
}
|
|
defer func() { _ = ch.Close() }()
|
|
if err := ch.InitSchema(cmd.Context()); err != nil {
|
|
return err
|
|
}
|
|
_, _ = fmt.Fprintf(cmd.OutOrStdout(), "schema ready: %s.%s\n", cfg.Index.Database, cfg.Index.Table)
|
|
return nil
|
|
},
|
|
}
|
|
cmd.Flags().BoolVar(&printOnly, "print", false, "print the DDL instead of executing it")
|
|
return cmd
|
|
}
|