Files
node-lookup/AGENTS.md
T
Ben Vincent 17ded87439
ci/woodpecker/pr/build Pipeline was successful
ci/woodpecker/pr/test Pipeline was successful
ci/woodpecker/pr/pre-commit Pipeline was successful
pdbmux: ship as k8s container, drop per-VM systemd/RPM delivery
The estate direction is all-in-kubernetes, so pdbmux (a long-running daemon)
should run as an in-cluster service rather than an RPM-installed systemd unit
on each VM. The RPM is for workstation/VM CLI tools only; a daemon does not
belong there.

- Remove packaging/pdbmux.service and drop pdbmux (binary, systemd unit,
  completions) from the RPM/nfpm spec and build-rpm.sh.
- Keep pdbmux in the Makefile build and the test suite.
- Add Dockerfile.pdbmux building a static CGO_ENABLED=0 binary on distroless
  (mirrors encapi's image style).
- Add .woodpecker/docker.yaml to build+push git.unkin.net/unkin/pdbmux:<tag>
  on v* tags via the docker-buildx plugin (droneci/DRONECI_PASSWORD creds,
  same as encapi), with k8s resources set.
- Update README/AGENTS.md: deployment is k8s, config via PDBMUX_* env.
2026-07-24 23:09:46 +10:00

195 lines
9.8 KiB
Markdown

# AGENTS.md
## Project Overview
This repo ships four related Puppet tools in one RPM:
- **`node-lookup`** — queries the PuppetDB API to retrieve and filter node facts.
- **`pburl`** — prints the Puppetboard node-page URL for each host (reads hosts
from args or piped `node-lookup` output). Output: `<host> <url>`.
- **`pblastreport`** — prints each host's last Puppet report time and its
Puppetboard URL. Output: `<host>\t<time>\t<url>`. Supports `--relative`/`-r`
(relative age) and `--timezone`/`-z <IANA>` (default: local timezone).
- **`pdbmux`** — a long-running HTTP daemon that presents a single merged
PuppetDB v4 query surface over the old (Consul) and new (k8s) PuppetDBs during
the VM→k8s migration. Merges `/pdb/query/v4/{nodes,facts}`, transparently
proxies other v4 paths to the primary, and exposes `/healthz`. See README.md
for full config/merge semantics.
`node-lookup` is the module root; `pburl`, `pblastreport` and `pdbmux` live
under `cmd/`. The three CLI tools share the `internal/puppet` package (config,
PuppetDB `nodes` queries, Puppetboard URL construction, stdin host reading);
`pdbmux` is self-contained (its own config + HTTP server).
## Structure
```
main.go # node-lookup CLI source (module root, package main)
main_test.go # node-lookup unit tests (mock PuppetDB via httptest)
cmd/pburl/main.go # pburl CLI
cmd/pblastreport/main.go # pblastreport CLI (report.go: report-time formatting)
cmd/pdbmux/ # pdbmux daemon: main.go, config.go, merge.go, server.go
internal/puppet/ # shared: config, puppetdb nodes query, board URLs, stdin
go.mod # Go module (module name: node-lookup)
go.sum # dependency checksums
Makefile # build / test / lint / completions / rpm / version-bump targets
packaging/nfpm.yaml # nfpm spec (envsubst-templated) for the RPM (CLI tools only)
Dockerfile.pdbmux # container image for the k8s-only pdbmux daemon
scripts/build-rpm.sh # generates completions + packages the RPM with nfpm
.woodpecker/ # CI: build, test, pre-commit (PR) + release/docker (tag)
dist/ # build output: binaries, completions, RPM (not committed)
```
Every binary is a separate `main` package, so `make build` builds each with its
own `-o` (a single `go build ./...` can't emit multiple mains to one file).
## Build
```bash
make build # -> dist/node-lookup (CGO disabled, static)
# or directly:
go build -o node-lookup ./...
```
Requires Go 1.21+. Dependencies: `github.com/spf13/cobra` (CLI), `gopkg.in/yaml.v3` (Ansible output).
## Packaging (RPM)
```bash
make rpm # build the binary + package it into dist/*.rpm via nfpm
```
`scripts/build-rpm.sh` generates bash/zsh/fish completions from the built binary
and bundles them alongside `/usr/bin/node-lookup`. On a `v*` tag the release
pipeline builds the RPM and `PUT`s it to the artifactapi `rpm-internal` repo.
The RPM contains the workstation/VM CLI tools only (`node-lookup`, `pburl`,
`pblastreport`). `pdbmux` is a k8s-only daemon and is deliberately excluded from
the RPM — it is released as a container image
(`git.unkin.net/unkin/pdbmux:<tag>`, built by `.woodpecker/docker.yaml` from
`Dockerfile.pdbmux`) and deployed via `argocd-apps`. `make build` and
`go test ./...` still cover pdbmux.
## Shell completions
Cobra provides a `completion` subcommand:
```bash
node-lookup completion bash # or zsh / fish / powershell
```
The RPM installs completions to the standard system paths
(`/usr/share/bash-completion/completions/`, `/usr/share/zsh/site-functions/`,
`/usr/share/fish/vendor_completions.d/`), so they work automatically once
installed. To load ad-hoc in the current shell, e.g. zsh:
`source <(node-lookup completion zsh)`.
## Running the Tool
```bash
./node-lookup --help
./node-lookup -R # show all nodes with role fact
./node-lookup -n <hostname> # lookup a specific node
./node-lookup -F <fact_name> # filter by fact name
./node-lookup -jF ipaddress,enc_role # several facts at once (comma-separated)
./node-lookup -R -m <value> # exact value match (-m)
./node-lookup -R -pm <value> # partial/regex match (-p -m combined)
./node-lookup -R -im <value> # inverse exact match (-i -m combined)
./node-lookup -R -ipm <value> # inverse partial match (-i -p -m combined)
./node-lookup -R -p <value> # value may also be given positionally
./node-lookup -R -1 # node names only
./node-lookup -R -2 # values only
./node-lookup -R -C # count occurrences
./node-lookup -R -A # output as Ansible YAML inventory (queried facts become host vars)
./node-lookup -j # output as JSON { host → { fact → value } }
./node-lookup --url http://host:8080/... # override PuppetDB URL for this invocation
echo -e "node1\nnode2" | ./node-lookup -R # pipe node names via stdin
```
### Companion tools
```bash
node-lookup -R | pburl # <host> <puppetboard-url> per line
pburl host1 host2 # hosts as args instead of stdin
node-lookup -R | pblastreport # <host> <last-report-time> <url>
pblastreport -r host1 # relative age (e.g. "3h ago")
pblastreport -z Asia/Singapore host1 # render the time in a specific IANA tz
```
Both read hostnames from arguments or the first field of each piped line (so
any `node-lookup` output mode works), de-duplicate, and share `node-lookup`'s
config file / env vars. `pblastreport` reads `report_timestamp` from the
PuppetDB v4 `nodes` endpoint (derived from the configured facts URL).
## Configuration
Precedence (lowest → highest): **defaults < config file < env vars < `--url` flag**
### Config file
XDG location: `$XDG_CONFIG_HOME/node-lookup/config.yaml` (default: `~/.config/node-lookup/config.yaml`)
```yaml
puppetdb_url: http://puppetdbapi.service.consul:8080/pdb/query/v4/facts
role_fact: enc_role
puppetboard_url: https://puppetboard.k8s.syd1.au.unkin.net # used by pburl / pblastreport
```
Generate the default config file:
```bash
./node-lookup config init
```
Show the active configuration (after all overrides applied):
```bash
./node-lookup config show
```
### Environment variables
| Variable | Config key | Description |
|---|---|---|
| `NODE_LOOKUP_URL` | `puppetdb_url` | PuppetDB facts endpoint |
| `NODE_LOOKUP_ROLE_FACT` | `role_fact` | Fact name used by `-R` flag |
| `NODE_LOOKUP_PUPPETBOARD_URL` | `puppetboard_url` | Puppetboard base URL (pburl / pblastreport) |
### CLI flag
`--url <url>` overrides the PuppetDB URL for a single invocation (highest precedence).
## Code Patterns
- **`loadConfig()`**: reads config file → applies env vars → returns `config` struct. Called once at startup in `main()`.
- **`buildQuery()`**: returns a PuppetDB PQL-compatible JSON array string. Uses `roleFact` from config (not hardcoded). Match modifiers: `-p` (partial/regex, uses `~` op), `-i` (inverse, wraps with `not`), composable.
- **Multiple facts**: `-F` accepts a comma-separated list (`ipaddress,enc_role`). `splitFactNames()`/`nameFilter()` turn several names into an `or` over `["=","name",<n>]` clauses; JSON output keys each value by the fact's real name so all requested facts appear per host.
- **Match value / `matchValue()`**: the value to match comes from `-m/--match` or, if that is empty, an optional positional argument. The positional fallback exists because pflag does not attach a space-separated value to a string flag grouped with a bool flag, so in `-pm k8s` the `k8s` arrives as a positional. `-m` still wins when both are given.
- **`queryPuppetDB(url, query)`**: takes the URL as a parameter — never reads globals.
- **`processResults()`**: iterates facts, returns sorted `"certname value"` strings. JSON string values are unquoted; other JSON types rendered as compact JSON.
- **Output modes**: JSON (`-j`), count (`-C`), Ansible YAML (`-A`), node-only (`-1`), value-only (`-2`), default (node + value). `-j` and `-A` share `factsByHost()`, so both attach the queried fact(s) per host — as an object under the host (`-j`) or as inventory host vars (`-A`).
- **Stdin support**: `stdinReader()` reads node names from stdin only when it is a real pipe/redirect carrying data (and no `-n` given). Terminals, `/dev/null`, and empty/closed pipes fall through to a normal query — so running without a TTY (e.g. invoked by an agent or CI) behaves like an interactive run instead of consuming empty input.
- **SIGPIPE handling**: `signal.Ignore(syscall.SIGPIPE)` so pipes to `head` etc. work cleanly.
## CLI Framework
Uses [Cobra](https://github.com/spf13/cobra). Root command is the query command. `config` is a subcommand with `init` and `show` sub-subcommands.
## Testing
```bash
make test # go test -v -race ./...
```
`main_test.go` covers query construction (all `-m`/`-p`/`-i` combinations), value
rendering, result processing/counting, config precedence (defaults < file < env),
`writeDefaultConfig`, the `stdinReader` no-TTY behavior, and every `run()` output
mode (default, `-1`, `-2`, `-C`, `-j`, `-A`, `-a`). PuppetDB is stubbed with
`httptest` — no live Consul/PuppetDB access is required.
## Gotchas
- `-1`, `-2`, `-C`, and `-A` all require `-R` or `-F`; the tool exits with an error otherwise.
- `-C` (count) with stdin reads all lines as pre-fetched `"node value"` output for counting — it does **not** query PuppetDB per line.
- JSON output (`-j`) builds `{ hostname: { factname: value } }` keyed by each result's actual fact name (so `-F ipaddress,enc_role` yields both per host); it falls back to the `-F` value, the `role_fact` config value (if `-R`), or `"value"` only when a result carries no name.
- `config init` fails if the config file already exists (will not overwrite).