Add try/confirm safe-apply with out-of-process revert #17
Reference in New Issue
Block a user
Delete Branch "benvin/safe-apply"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
A firewall change that cuts off SSH leaves the host unreachable until someone reaches the console.
tomswall try: snapshot to/var/lib/tomswall, apply, revert unless confirmed in timerevert --id <try>) so the revert runs even iftryis killed; a stale timer cannot revert a newer tryconfirm(fails if already reverted) andrevert; a failed revert keeps the snapshot for retryapply,flush,purge, a secondtryand agent applies refuse while atryis pendingreadCurrentStatenow errors instead of returning empty stateRefs #11
kill -9/OOM/crash oftrymid-wait leaves the bad ruleset applied with nothing left to revert it; operator stays locked out (issue acceptance fails) → arm an out-of-process fallback before applying (e.g.systemd-run --on-active=<timeout> tomswall revert-pendingusing a snapshot persisted under /run/tomswall, cancelled on confirm).confirmprints "Confirmed." as soon as SIGUSR1 is sent; if the timeout/abort already fired the signal is swallowed, revert runs, and the operator is told it was confirmed → maketryack (or confirm wait for try's exit and report its result); do not report success from a fire-and-forget signal.tryoverwrites it, snapshots the first try's unconfirmed ruleset (so its revert restores the bad state), and the first'sdefer os.Removedeletes the second's pidfile → flock the pidfile, refuse if held.errors.As(err, *url.Error)matches ctx cancellation (agent shutdown), TLS/cert errors and single transient failures; a good config is reverted and its generation is skipped → checkctx.Err(), treat only network/dial/timeout errors as unreachable, retry a few times before reverting.status: reverted"; nothing is reported after the revert → report reverted status (generation) once the API is reachable again.revertedis in-memory only; agent restart re-applies the bad generation → persist it (next to the cache).tryboth mutate the table with no coordination; an agent cycle during the try window is clobbered by the try revert (or vice versa) → agent skips apply while the try lock is held.ensureChains, which resets chain policies to the hard-coded drop/accept; policies are not part of the snapshot → capture and restore chain policies.readCurrentStatenow returns errors instead of an empty state; this changesplan/apply/statusbehaviour outside this PR's why → keep, but call out in the body, or split into its own PR.confirm(stale pidfile, race with timeout), signal wiring, or the table-absent Snapshot/Restore path (engine.go:206) → add tests for confirm and for the nil-snapshot restore.selectpicks randomly → prefer abort/confirm deterministically.Add safe-apply: try/confirm and agent auto-revertto Add try/confirm safe-apply with out-of-process reverttomswall applynever callstryapply.Acquire, so it runs during a pending try and the later confirm/timeout revert silently overwrites it with the stale snapshot → take the lock and returnErrPendingthere as the agent does.Revert()with a pending snapshot is never exercised (only the nothing-pending path); restore ordering,Discardafter restore, and revert-failure-keeps-snapshot are untested → add a test that injects an engine/restore seam and asserts snapshot removal and timer stop on success, retention on error.Restoreis tested; the present-table path (decodeState+applywith policies) has no engine-level test → add one asserting the sent del/add rule messages and chain policies.AcquirereturnsErrPendingforever → documenttomswall revertas the recovery in the error text.No findings.