# tomswall control plane — design Status: draft / theory-craft. No implementation yet. This document specifies a **fleet control plane** for tomswall: a central, Terraform-managed API that lets you declare zones, address groups, and firewall policy **once** and have every device in a routed estate enforce a provably consistent posture. tomswall itself remains the per-host nftables engine; this adds an orchestration layer above the fleet. --- ## 1. Motivation Today each tomswall host owns a complete, independent `tomswall.yaml` — its own zones, interfaces, policy, and rules. In an estate of many firewalls and routers that is unmanageable: a single logical intent ("hosts in zone A may reach zone B on tcp/22") has to be hand-translated into per-hop rules on every device along every possible path. The control plane inverts this. You declare the **intent once**; the API compiles it into the concrete rules each device needs and serves each device its rendered config. A connection that crosses several firewalls — `src → rt1 → rt2 → rt3 → dest` — is expressed as a single rule. Key environment facts this design is built around: - Internal routing is **dynamic (FRR)** with **ECMP** — paths are not stable and cannot be pinned. - **NAT happens only at the edge**; the internal core is purely routed. - Devices fall into two classes: **routers** (FRR core) and **firewalls** (zone boundary enforcement). --- ## 2. Terminology | term | meaning | |---|---| | **zone** | A fleet-global named network segment (a set of subnets). Replaces per-host zones. | | **subzone** | A zone nested inside a parent zone; its subnets must be ⊂ the parent's. | | **address group** | A fleet-global named set of addresses → materialized as an nftables named set (ipset). Population source is `static`, `dns`, or `asn`. | | **fabric** | A routing domain: a group of routers (and the firewall edges attached to it) that share an FRR routing fabric. | | **binding** | A per-device mapping of a global zone to that device's local interface(s). The only legitimately host-specific object. | | **intent / rule** | A fleet-global `source → dest` policy statement, matched on zones and/or address groups. | | **generation** | A monotonic version stamp on rendered config; devices report the generation they have applied. | --- ## 3. Architecture ``` terraform ──▶ tomswallapi ◀── peers with FRR (BGP-LS / BMP) │ · inventory + topology (devices, zones, fabrics, bindings) │ · match/address-group model (static/dns/asn ipsets) │ · compiler: intents → per-device tomswall.yaml │ · reachability validation from routing state │ · per-device config + set-delta feed (authenticated) ▼ fw-a rt1 rt2 rt3 fw-b ── each runs `tomswall agent`: pull config → differential apply → report generation + on-device resolver maintains dns ipsets ``` Three moving parts: 1. **`tomswallapi`** — the control plane. Stores the model (Postgres, per the house stack), peers with FRR for routing/reachability, compiles intents into per-device configs, and serves them. 2. **`tomswall agent`** — a new pull mode on the existing binary. Periodically fetches its rendered config, runs the existing differential `apply`, maintains dns-backed ipsets via an on-device resolver, and reports its applied generation. 3. **FRR peering** — the API consumes routing state (BGP-LS for topology, BMP / route monitoring for prefix reachability) so it can validate that a zone's subnet really originates where its firewall claims, and scope which routers an intent must touch. --- ## 4. Data model: two layers The founding constraint: > Every host may have a different zone→interface mapping, but all share the same > zones, address groups, policies, and rules. That splits the model into two layers with a hard wall between them. ### 4.1 Global layer (defined once, byte-identical everywhere) `zones` · `subzones` · `address_groups` · `portgroups` · `policies` · `rules` · `blrules` · `conntrack` · `secmarks` · `vars` · `fabrics`. **None of these ever names an interface.** The rule `zone-a → zone-b tcp/22` is the same object on every device. ### 4.2 Device layer (the only per-host freedom) `class` (router/firewall) · `fabric` membership · per-device `settings` overrides · `resolver` config · and the **binding table** (`zone → interface(s)`). The rendered per-device `tomswall.yaml` = **global rules projected through that device's binding table**. Same intent, different local interface resolution: ``` GLOBAL (shared): rule zone-a → zone-b tcp/22 accept zones zone-a=10.1.0.0/24 zone-b=10.4.0.0/24 fw-a bindings: zone-a → eth1 ; → bond0 (fabric "core") fw-b bindings: zone-b → ens5 ; → ens4 (fabric "core") rt3: no zone bindings ; fabric "core" ``` `zone-a` is *attached* on fw-a (resolved via its eth1 binding) but *remote* on fw-b and rt3 (resolved purely by subnet match on the core side). That asymmetry is fine because the compiled rule is interface-agnostic (§7). --- ## 5. Device classes and fabrics - **firewall** — a zone boundary. Zones "live" on firewall interfaces; a firewall enforces with default-drop between zones and is where the real policy edge sits. - **router** — an FRR core member belonging to a **fabric**. Enforces interface-agnostically (any interface, address-matched) because ECMP means the path is not pinnable. A **fabric** is a routing domain with an `enforce_on_routers` flag: - `enforce_on_routers = false` (transparent transit) — routers route freely for internal ranges and rely on conntrack; only boundary firewalls enforce the intent. Fewest rules; no in-core containment. - `enforce_on_routers = true` (defense-in-depth) — every router in the fabric also carries the intent as an interface-agnostic permit with default-drop transit. Every hop is a checkpoint; contains a compromised core node. Per-fabric choice lets a small trusted core run transparent while a larger or multi-tenant core runs defense-in-depth. --- ## 6. The match model Source and dest use a **shorewall-style** grammar. Each direction is a comma-list of elements; **within an element a zone gates the selector (AND); across elements the list is a union (OR)**. ``` element := zone # bare zone — the zone-to-zone base case | zone ":" selector # zone AND selector (selector must be paired) selector := "+" ipset # named address group (static / dns / asn) | "&" fqdn-group # (fqdn groups are just dns-typed ipsets; "+" also accepted) source = "loc, net:+asn_cloudflare, dmz:+partner_api" # loc OR (net AND asn_cloudflare) OR (dmz AND partner_api) ``` Rules: - **Bare zone → legal.** `loc → net`, the zone-to-zone base case. - **`zone:+selector` → legal.** The selector must always be paired with a zone. - **Bare selector → rejected at plan time.** No floating `+ipset` / `asn:` without a zone. Why the pairing is structural, not cosmetic: an internet-facing zone like `net` has no finite subnet — it is "everything else" — so it can never stand as a clean address match on its own. Pairing supplies the missing halves: **the zone gives direction/interface, the selector gives concrete addresses.** ``` rule: loc → net:+asn_cloudflare tcp/443 edge fw render: oif= daddr @asn_cloudflare tcp dport 443 accept # zone `net` → the internet-facing binding; # asn_cloudflare → the actual prefixes to match ``` Internal zones (which have subnets) may stand bare; internet/edge zones effectively require a selector. --- ## 7. Compilation ### 7.1 Interface-agnostic, address-matched rules (the ECMP unlock) Because FRR picks paths dynamically and load-balances across ECMP, rules must **not** be compiled to per-hop `iif/oif`. Every device carries the same rule matched on `saddr ∈ source, daddr ∈ dest, proto, port` in the forward chain, on **any interface**: ``` rule: zone-a (10.1.0.0/24 @ fw-a) → zone-b (10.4.0.0/24 @ fw-b) tcp/22 fw-a (firewall): forward saddr 10.1.0.0/24 daddr 10.4.0.0/24 tcp dport 22 accept rt1…N (routers): forward saddr 10.1.0.0/24 daddr 10.4.0.0/24 tcp dport 22 accept fw-b (firewall): forward saddr 10.1.0.0/24 daddr 10.4.0.0/24 tcp dport 22 accept ``` - **Return traffic needs no rules.** Each device is independently stateful, so every hop's own `ct state established,related accept` handles the reply. Only the forward direction is emitted. - **ECMP and asymmetric routing just work** — flow #1 may go rt1→rt3, flow #2 rt1→rt2→rt3, and the return may take a third path; every device it touches already permits the 5-tuple and holds its own conntrack state. ### 7.2 Over-approximation is safe → no exact path computation Because rules are interface-agnostic and stateful, programming a permit on a router the traffic never crosses just leaves an unused rule. So the compiler never needs FRR's *actual* chosen path — only **which fabric(s) could carry A↔B**, which is coarse and static-friendly. FRR peering (§3) is used to *narrow* the router set (and to validate zone origins), not to pin a path. ### 7.3 nftables set / map form for scale Every zone and address group is a **named nft set** (`flags interval` for CIDRs). Rules reference sets by name. Where many intents land on the same core routers, compile into **sets / verdict maps** keyed on `(saddr, daddr, proto, dport)` rather than long linear chains, so match cost stays sub-linear. `zone:+ipset` compiles to a conjunction of two set lookups (`ip saddr @z_zone` **and** `ip saddr @g_ipset`); a multi-element list becomes multiple rule lines sharing a verdict. The critical property: **membership is the only thing that churns; rules are stable.** Adding/removing an address is a set-element op — no rule reload, no chain rebuild, existing connections preserved. --- ## 8. Address groups (ipsets) An address group is always an nftables named set. What differs is only how its members are populated: | type | member source | resolved where | refresh | |---|---|---|---| | `static` | explicit CIDRs/IPs | API (constant) | n/a | | `dns` | FQDNs → A/AAAA | **on-device resolver** | per record TTL | | `asn` | ASN(s) → prefixes via iplocate | **central API** | configurable, default 24h | ### 8.1 ASN groups An ASN group is defined **once, globally**, and materializes a set named by convention `asn_` (friendly) or `asn_`. An ISP may own several ASNs; one group merges them into one set. Rules reference it like any ipset — `net:+asn_cloudflare` — there is **no** special `asn:` selector token. ```hcl resource "tomswall_address_group" "asn_cloudflare" { type = "asn" asns = ["13335", "209242"] refresh = "24h" # default; configurable per group / globally } # → materializes nft set asn_cloudflare ``` Expansion is **centralized**: the API holds a single iplocate key (in Vault), calls the ASN data-type endpoint, caches prefixes per ASN, refreshes on the TTL, and ships prefix deltas as set-element ops. Devices never call iplocate. ASN membership is therefore **fleet-uniform**. ### 8.2 DNS groups and the per-host resolver DNS groups resolve **on-device**, so each site honors its own split-horizon / geo-DNS view. The resolver used follows a precedence chain: ``` per-device resolver override → fleet default_resolver → system /etc/resolv.conf ``` ```hcl resource "tomswall_device" "fw_a" { resolver = ["10.1.0.53", "10.1.0.54"] } # global: default_resolver = ["10.0.0.53"] (or "system") ``` A device in site A resolves `vpn.corp` against site A's resolver and populates *its own* set from that view; site B may legitimately differ. DNS membership is **not** guaranteed fleet-uniform — only the rules and set names are. Record TTL is used as the nft **element timeout**, so stale entries self-evict and the resolver refreshes before expiry. ### 8.3 Set lifecycle notes - v4 and v6 members are held in parallel family sets (or inet-family sets). - Element deltas (`nft add/delete element`) are a separate, lighter feed than rule rollout; they never bump the rule generation. --- ## 9. NAT / masquerade / netmap / policy — full parity The control plane is a **superset** of the tomswall config model, never a lossy subset: the rendered per-device output *is* a full `tomswall.yaml`, so anything tomswall can express is reachable. **Every section is a typed, first-class resource — there is no raw-YAML escape hatch.** Sections are handled in one of three tiers: - **Global-compiled** — defined once, projected identically to every relevant device: `policy` (default posture), `rules`, `portgroups`, `zones`/subzones, `blrules`, `conntrack`, `secmarks`, `vars`. - **Global-intent, device-resolved** — defined once against zones; each device resolves to its own interfaces via its binding table: `snat`/**masquerade**, `netmap`, `nat` (1:1). Example: `masquerade source=loc egress=net` lands only on devices that bind **both** `loc` and `net` (i.e. edge firewalls); interior routers have no `net` binding and skip it automatically. - **Per-device** — declared against a device (or a selector): `host`, `provider`, `rtrules`, `route`, `tunnel`, `stopped_rule`, `proxy_arp`/`ndp`, `arp_rule`, `maclist`, `accounting`, `mangle`, `tc_*`. --- ## 10. Terraform resource catalog ### Fleet / topology - `tomswall_device` — name, `class` (router/firewall), `fabric`, per-device `settings` overrides, `resolver` - `tomswall_fabric` — routing domain, `enforce_on_routers` - `tomswall_binding` — zone→interface map (per device+zone) - `tomswall_settings` — global defaults (address_family, log_level, ip_forwarding, table_name) ### Global-compiled - `tomswall_zone` — subnets, `parent` (subzones) - `tomswall_address_group` — ipset; `type` = static/dns/asn - `tomswall_portgroup` - `tomswall_policy` — default zone→zone posture, `priority` - `tomswall_rule` — the `zone:+ipset` / `&fqdn` intents - `tomswall_blrule` - `tomswall_conntrack` - `tomswall_secmark` - `tomswall_var` ### Global-intent, device-resolved - `tomswall_snat` — masquerade/SNAT (auto-scopes to devices binding both zones) - `tomswall_netmap` — anchored subnet↔subnet - `tomswall_nat` — 1:1 static, bound to the device holding the public IP ### Per-device (`device` reference or a `class`/`fabric`/`all` selector) - `tomswall_host` - `tomswall_provider` - `tomswall_routing_rule` (rtrules) - `tomswall_route` - `tomswall_tunnel` - `tomswall_stopped_rule` - `tomswall_proxy_arp` / `tomswall_proxy_ndp` - `tomswall_arp_rule` - `tomswall_maclist` - `tomswall_accounting` - `tomswall_mangle` - `tomswall_tc_device` / `tomswall_tc_class` / `tomswall_tc_filter` / `tomswall_tc_interface` / `tomswall_tc_priority` ### Data sources - `tomswall_device_config` — rendered `tomswall.yaml` preview for a device - rule fanout preview — which devices an intent will touch (surfaced in `plan`) Per-device resources accept **either a single `device` or a selector** so common objects (a shared static route, a provider) are declared once, preserving the define-once ethos even in the local tier. --- ## 11. Agent protocol `tomswall agent` (or a systemd timer invoking a pull) does: 1. **Pull** its rendered config by `device_id` from the API (authenticated). 2. Write it to a **local cache file**. 3. Run the existing **differential `apply`** (compute diff vs live nftables, apply only the delta atomically; never tears the firewall down). 4. Maintain **dns ipsets** via the on-device resolver (add/delete elements on TTL). 5. **Report** the applied `generation` back to the API. ### 11.1 Do not fail closed On API-unreachable: **keep the cache, re-apply it (idempotent no-op), never flush to deny.** Existing rules ride through control-plane outages untouched; only *changes* require the API. This is a deliberate availability choice. ### 11.2 Rollout & convergence A rule spanning several devices rolls out as each device pulls independently. Mid-rollout the connection is blocked at whichever hop has not yet pulled — i.e. **fail-closed for adds** (safe). Config is **generation-stamped** and devices report the generation applied, giving a fleet-wide "converged / N behind" view. ### 11.3 tomswall engine changes required - A **pull/agent mode** (`tomswall agent`, or `source: http` in the config loader). - The **interface-agnostic rule form** (match `saddr`/`daddr` with no `iif/oif`). - **Resolver config** and on-device dns-set maintenance. - **Ownership tags** (resource-id + generation as nft comments) so `purge` / foreign-rule detection never fights control-plane-managed content. --- ## 12. Invariants (validated at plan time) 1. **Zones/policies/rules/groups are global-only.** A device may *bind* a zone to an interface; it may never *define* one. No local zone namespace. 2. **Subnet→zone is globally unique.** Every subnet belongs to exactly one zone fleet-wide; no overlaps. (Ambiguous saddr/daddr matching otherwise.) 3. **Every zone in a rule must be resolvable on every enforcing device** — either locally bound (attached) or remote-reachable via its fabric. Unresolvable → rejection, not a silent gap. FRR reachability proves this. 4. **Each zone is anchored at ≥1 firewall interface** — the API knows where every zone physically enters the fabric, cross-checked against FRR-advertised origins. 5. **A firewall must bind every directly-connected zone.** A connected subnet with no zone identity is rejected/flagged (checked against FRR-reported prefixes). Zones the device does not attach are implicitly remote — no verbose "not here" declarations needed. 6. **Subzones nest by containment** — a subzone's subnet ⊂ its parent's; hierarchy is global. 7. **Interface names appear only in the binding table.** An interface name leaking into any global object → rejection. 8. **A selector never appears without a zone.** `zone:+ipset` / `zone:&fqdn` only; bare `+ipset` / `asn:` → rejection. Bare zones remain legal. 9. **The zone in a pair supplies direction/interface; the selector supplies addresses** — this is what lets no-subnet zones (`net`, edge) participate. 10. **Every tomswall section is a typed resource.** No raw-YAML passthrough; nothing bypasses validation. 11. **Config is generation-stamped**; devices report the generation applied. --- ## 13. Fail-safe semantics - **Resolution failure keeps last-good membership.** An iplocate outage or DNS SERVFAIL must never empty a set. (Matches the agent's don't-fail-closed stance.) - **A genuinely-empty group (NXDOMAIN, ASN with no prefixes) makes its rule inert and is logged/flagged** — never "match everything." An unresolvable source/dest disables its rule loudly, never opens it. - **Adds fail closed, the control plane fails open.** Partial rollout blocks new flows until every hop converges; a dead API leaves the last-good posture running. --- ## 14. Security / auth - Agents authenticate to the API (mTLS or Vault-issued per-device tokens, per the house pattern). - The iplocate API key and any resolver credentials live in Vault. - The API is the single source of truth; state in Postgres. - Compiled objects carry ownership tags so the on-device engine can distinguish control-plane content from local/foreign rules. --- ## 15. Open questions / future work - **FRR integration depth** — BGP-LS (topology) vs BMP / route monitoring (prefix reachability) vs a lighter agent-reported FIB. Start with what proves zone origin and fabric membership; deepen as needed. - **Set-union match ergonomics** — whether multi-element source/dest compiles to multiple rule lines or a merged interval set; membership churn is handled at the member-set level regardless. - **NAT along non-edge paths** — out of scope by assumption (routed core, edge-only NAT). Revisit only if internal translation is ever introduced. - **Multi-tenancy** — whether fabrics/zones need tenant scoping for RBAC.