e24c35f534
## Why Builds on #107 (merged), which derives `github_rpm` RPM metadata lazily on the client request path, single-flighted per replica. Two problems remain: the derive still happens per replica, so across a multi-replica deployment the same releases are scanned and re-derived N times, multiplying GitHub queries; and a cold cache blocks the first request on a full derive. GitHub's rate limits are low (~60/hr unauthenticated, ~5000/hr authenticated), so this needs a single coordinated syncer with a shared rate limit and conditional requests. ## How - Add a single per-process background syncer (started at boot, cleanly stopped on shutdown) that owns a deduped/coalescing work queue, a worker pool, and one global token-bucket rate limiter (`golang.org/x/time/rate`) bound onto the github provider so every GitHub call (releases list + each ranged asset GET) acquires a token first. - Re-check each `github_rpm` remote for new/changed releases on its existing `mutable_ttl` cadence; derive only new/changed assets incrementally and prune assets that disappear upstream. Repodata is served from primed DB rows. - Prime metadata in the background on remote creation; the create call returns immediately. - Send the stored releases-list `ETag` as `If-None-Match`; a `304` derives nothing and is not counted against GitHub's rate limit, so an unchanged repo is nearly free. - Coordinate replicas through a `github_rpm_sync_state` row (`last_synced_at`, `etag`, `sync_lease_owner`, `sync_lease_expires`): a periodic scan runs only for the replica that atomically claims the lease, bounding total GitHub load to ~once per `mutable_ttl` regardless of replica count; the ETag is shared through the same row. - Keep the request path fast: serve current cache, enqueue a prime on an empty cache, and return a bounded wait then a retryable `503` rather than blocking on a cold derive. - Add `GITHUB_SYNC_RATE` / `GITHUB_SYNC_BURST` / `GITHUB_SYNC_WORKERS` / `GITHUB_SYNC_POLL_INTERVAL` config with conservative defaults (1 req/s, burst 5, 3 workers, 60s tick) and document the syncer in the README. ## Tests - Unit (httptest, Range/ETag-aware fixture): `304` releases response derives nothing; incremental derive fetches only the newly added asset; the shared limiter caps request rate; work-queue enqueues coalesce to one job; prime enqueues a job; a held lease stops a second replica from scanning; cold-start serves `503` while warm cache serves `200`. - DB integration (testcontainers postgres): the real lease SQL — one holder at a time, recency gate blocks a too-soon periodic re-claim, prime (freshness 0) bypasses recency but respects a live lease. - Docker e2e re-run: `dnf install dotvault` works; prime-on-create derives in the background at ~1 req/s (global limiter); `dnf makecache` served fast from the priming cache (no cold block); clean shutdown mid-scan, no panics. ## Notes - Reuses `mutable_ttl` as the check interval (no new per-remote field), per brief. Reviewed-on: #108 Co-authored-by: Ben Vincent <ben@unkin.net> Co-committed-by: Ben Vincent <ben@unkin.net>
74 lines
2.8 KiB
Go
74 lines
2.8 KiB
Go
package database
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"time"
|
|
|
|
"github.com/jackc/pgx/v5"
|
|
|
|
"git.unkin.net/unkin/artifactapi/pkg/models"
|
|
)
|
|
|
|
// ListGitHubRPMRemotes returns every github_rpm remote so the syncer can sweep
|
|
// them on each poll tick.
|
|
func (db *DB) ListGitHubRPMRemotes(ctx context.Context) ([]models.Remote, error) {
|
|
rows, err := db.Pool.Query(ctx, `SELECT `+remoteCols+` FROM remotes WHERE package_type = $1 ORDER BY name`, models.PackageGitHubRPM)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
defer rows.Close()
|
|
|
|
var remotes []models.Remote
|
|
for rows.Next() {
|
|
var r models.Remote
|
|
if err := scanRemote(rows, &r); err != nil {
|
|
return nil, err
|
|
}
|
|
remotes = append(remotes, r)
|
|
}
|
|
return remotes, rows.Err()
|
|
}
|
|
|
|
// ClaimGitHubSyncLease atomically claims the per-remote sync lease. It succeeds
|
|
// (claimed=true) only when the remote is due — never synced, or synced longer
|
|
// than freshness ago — and no live lease is held by another replica. This bounds
|
|
// total GitHub load to roughly one scan per freshness window regardless of how
|
|
// many replicas poll. The returned etag is the stored releases-list ETag, shared
|
|
// across replicas so a conditional request can short-circuit an unchanged repo.
|
|
// A zero freshness (used for prime scans) ignores the recency gate and claims
|
|
// whenever no live lease is held.
|
|
func (db *DB) ClaimGitHubSyncLease(ctx context.Context, remoteName, owner string, freshness, lease time.Duration) (bool, string, error) {
|
|
row := db.Pool.QueryRow(ctx, `
|
|
INSERT INTO github_rpm_sync_state AS s (remote_name, sync_lease_owner, sync_lease_expires)
|
|
VALUES ($1, $2, now() + make_interval(secs => $4))
|
|
ON CONFLICT (remote_name) DO UPDATE
|
|
SET sync_lease_owner = $2,
|
|
sync_lease_expires = now() + make_interval(secs => $4)
|
|
WHERE (s.last_synced_at IS NULL OR s.last_synced_at < now() - make_interval(secs => $3))
|
|
AND (s.sync_lease_expires IS NULL OR s.sync_lease_expires < now())
|
|
RETURNING s.etag
|
|
`, remoteName, owner, freshness.Seconds(), lease.Seconds())
|
|
|
|
var etag string
|
|
if err := row.Scan(&etag); err != nil {
|
|
if errors.Is(err, pgx.ErrNoRows) {
|
|
return false, "", nil
|
|
}
|
|
return false, "", err
|
|
}
|
|
return true, etag, nil
|
|
}
|
|
|
|
// ReleaseGitHubSyncLease records the completed scan and frees the lease. Only the
|
|
// owning replica may release; last_synced_at advances so the next poll waits a
|
|
// full freshness window, and etag is persisted for the next conditional request.
|
|
func (db *DB) ReleaseGitHubSyncLease(ctx context.Context, remoteName, owner, etag string, syncedAt time.Time) error {
|
|
_, err := db.Pool.Exec(ctx, `
|
|
UPDATE github_rpm_sync_state
|
|
SET last_synced_at = $3, etag = $4, sync_lease_owner = '', sync_lease_expires = NULL
|
|
WHERE remote_name = $1 AND sync_lease_owner = $2
|
|
`, remoteName, owner, syncedAt, etag)
|
|
return err
|
|
}
|