From iptables to nftables: a practical migration
Translate your ruleset to nftables' cleaner syntax and sets, and run both during the cutover without lockouts.
Most Linux hosts have already moved to nftables without anyone deciding to: on current Debian, Ubuntu, RHEL and Fedora the iptables command is a compatibility shim (iptables-nft) that writes nftables rules through the old syntax. The shim keeps scripts running. What it does not give you is one ruleset for both address families, named sets instead of ipset, or an atomic load, and it leaves you debugging two layers of abstraction when a rule does not do what you expect. Migrating means writing the ruleset in native nft syntax once, and doing the cutover in a way that cannot lock you out of SSH.
A host firewall on a cloud VM is the second gate behind the security group, not the first, so the cost of a mistake is usually lost access rather than exposure. That changes the design priority: the migration below is ordered around never applying a policy drop chain without a working allow rule in front of it.
Start from a translation, not from a blank file
iptables-save is the inventory. iptables-restore-translate -f converts that whole dump into nftables statements, one chain per legacy chain with the same names and priority 0, which is a faithful but ugly first draft. iptables-translate does the same for a single command and is useful for checking how one match maps. Neither produces sets, the inet family, or sane chain names, so the draft is what you rewrite from, not what you deploy.
iptables-save > /root/iptables-backup.rulesiptables-restore-translate -f /root/iptables-backup.rules > /root/draft.nftiptables-translate -A INPUT -p tcp --dport 22 -m conntrack --ctstate NEW -j ACCEPTnft add rule ip filter INPUT tcp dport 22 ct state new counter acceptiptables -Viptables v1.8.10 (nf_tables)(nf_tables) means the shim is already in use; (legacy) means the old backend is loaded tooThe rewritten ruleset
Three changes turn the draft into something maintainable. The inet family gives one table that sees IPv4 and IPv6, so the ip6tables copy of every rule disappears. A named set replaces the list of admin addresses that would otherwise be repeated per rule and can be updated at runtime with nft add element. And the file starts with flush ruleset, which makes loading it atomic: nft -f applies the whole file in one transaction, so the host is never in the half-configured state a shell script of iptables -A lines passes through.
#!/usr/sbin/nft -fflush rulesettable inet filter {set admin_v4 {type ipv4_addrelements = { 203.0.113.10, 203.0.113.11 }}chain input {type filter hook input priority filter; policy drop;ct state established,related acceptct state invalid dropiif lo acceptip protocol icmp acceptip6 nexthdr icmpv6 accepttcp dport 22 ip saddr @admin_v4 accepttcp dport { 80, 443 } acceptlimit rate 5/minute log prefix "nft-drop: "counter drop}chain forward {type filter hook forward priority filter; policy drop;# container hosts: Docker's own FORWARD accept does not carry over to this# chain (see priorities below), so published ports need an accept here tooct state established,related acceptiifname "docker0" acceptoifname "docker0" ct state new accept}}
Two lines are there because their absence is a classic outage. ct state invalid drop before the accept rules stops malformed packets from being treated as part of an established flow. ip6 nexthdr icmpv6 accept is not optional on a dual-stack host: neighbour discovery is ICMPv6, and dropping it takes IPv6 connectivity down within minutes while everything looks fine over IPv4.
# order-dependent, IPv4 only, ip6tables copy maintained separatelyiptables -A INPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPTiptables -A INPUT -i lo -j ACCEPTiptables -A INPUT -p tcp -s 203.0.113.10 --dport 22 -j ACCEPTiptables -A INPUT -p tcp -s 203.0.113.11 --dport 22 -j ACCEPTiptables -A INPUT -p tcp -m multiport --dports 80,443 -j ACCEPTiptables -P INPUT DROP
Priorities: where your chain runs relative to Docker and conntrack
Within one hook, Netfilter runs base chains in increasing numeric priority, and nftables predefines no chains at all: every base chain you want is one you create. The values matter as soon as something else on the host registers chains too. Docker publishes ports through a dstnat chain at -100 in prerouting, so by the time filter chains run the destination is already the container and the packet traverses forward, never input. On forward, verdicts combine in a way that surprises people coming from iptables: accept ends only the base chain that issued it and the packet continues to the next base chain on that hook, whereas drop anywhere is final. A chain forward { policy drop; } of your own therefore blocks published container ports even though Docker's FORWARD chain accepted them, so the ruleset above accepts bridge traffic explicitly. Reading nft list ruleset after the container runtime has started is the only reliable way to see every chain on a hook.
Standard priorities within a hook (inet, ip, ip6)
| Name | Value | Runs | Used by |
|---|---|---|---|
raw | -300 | before conntrack | notrack rules |
mangle | -150 | after conntrack | marking, TTL changes |
dstnat | -100 | prerouting | Docker port publishing, kube-proxy |
filter | 0 | the usual place for accept/drop | your input and forward chains |
security | 50 | after filter | SELinux secmark |
srcnat | 100 | postrouting | masquerade for containers |
The cutover order
nft -c -f parses and validates the file without touching the kernel, which catches syntax errors but not logic errors: a valid file can still drop your session. So the load happens from a session that has a way back. Keep a second SSH session open, or better, have console access, and before loading a policy drop ruleset schedule a rollback that runs unless you cancel it. Then load, test from the second session, and cancel the rollback.
nft -c -f /etc/nftables.conf && echo syntax oksyntax oknft list ruleset > /root/before.nftsystemd-run --on-active=120 --unit=nft-rollback nft -f /root/before.nftin 120 s the previous ruleset comes back unless this timer is stoppednft -f /etc/nftables.confssh -o ConnectTimeout=5 admin@203.0.113.50 true && echo "new session works"new session workssystemctl stop nft-rollback.timersystemctl enable nftablesThe nftables service is the persistence step: on Debian and Ubuntu it loads /etc/nftables.conf at boot, while RHEL and Fedora read /etc/sysconfig/nftables.conf, which includes files from /etc/nftables/. If the host also runs Docker, restart the runtime after the first load and check that its chains came back; flush ruleset removes them too, which is a second reason the container runtime should start after the firewall rather than before it.
Once the input chain is stable, the same file is the place for an output chain that limits what a compromised process can reach, with explicit exceptions for DNS, NTP and the package mirror. The drop log line above already produces the events; Linux detection engineering covers turning nft-drop: lines and rule changes into alerts, and Linux hardening covers the SSH and systemd settings the firewall sits in front of.