Writing an nftables ruleset

The ruleset underneath, loaded atomically.

Intermediate16 min · lesson 17 of 24

Under ufw and firewalld sits nftables, the kernel's packet filter, and on a host without a front-end you can write its ruleset yourself. That gives you everything the front-ends hide: named sets you can change without a reload, counters on each rule, logging exactly where you want it, and one file that is the complete firewall. In this lesson you write a default-deny ruleset for a web server, check and load it atomically, prove it from two other hosts, change it live, arm a rollback that really runs when you lock yourself out, and make it load at boot on Ubuntu and RHEL. Pick this route instead of a front-end, never alongside one: the host firewall lesson showed what happens when two managers share a ruleset.

The ruleset

As in the host firewall lesson, the lab's server is a network namespace, hard-nft-web: a separate copy of the kernel's network stack with its own interfaces, ports and ruleset, so nothing here can cut off the lab VM. It sits between an admin host (192.0.2.10) and an outside host (203.0.113.50), both namespaces too, joined to it by veth pairs (virtual cables). This script builds all three and starts three listeners on the server with socat (sudo apt install socat on a default install): SSH on 22, a web server on 80 and a metrics exporter on 9100.

/var/tmp/hard-nft-lab.sh
#!/bin/sh
# Lab network for the nftables lesson, built from network namespaces: a server
# (hard-nft-web) between an admin host (hard-nft-adm, 192.0.2.10) and an
# outside host (hard-nft-out, 203.0.113.50). Needs socat: apt install socat.
set -e
for n in hard-nft-web hard-nft-adm hard-nft-out; do ip netns add $n; ip -n $n link set lo up; done
ip link add nft-web-a netns hard-nft-web type veth peer name nft-adm netns hard-nft-adm
ip link add nft-web-b netns hard-nft-web type veth peer name nft-out netns hard-nft-out
ip -n hard-nft-web addr add 192.0.2.1/24 dev nft-web-a
ip -n hard-nft-web addr add 203.0.113.1/24 dev nft-web-b
ip -n hard-nft-adm addr add 192.0.2.10/24 dev nft-adm
ip -n hard-nft-out addr add 203.0.113.50/24 dev nft-out
ip -n hard-nft-web link set nft-web-a up
ip -n hard-nft-web link set nft-web-b up
ip -n hard-nft-adm link set nft-adm up
ip -n hard-nft-out link set nft-out up
# Listeners on the server: SSH (22), web (80) and a metrics exporter (9100).
for p in 22 80 9100; do
systemd-run --quiet --unit=hard-nft-l$p -p NetworkNamespacePath=/run/netns/hard-nft-web \
socat TCP-LISTEN:$p,fork,reuseaddr SYSTEM:"echo port-$p"
done
deploy@web01 · Ubuntu 26.04 LTS
$ sudo /var/tmp/hard-nft-lab.sh ip netns list
hard-nft-out hard-nft-adm hard-nft-web

Commands for the server start with sudo ip netns exec hard-nft-web; on a real server you type what follows. ip netns exec also bind-mounts files from /etc/netns/hard-nft-web/ over their namesakes in /etc, so inside the namespace /etc/nftables.conf is the lab's own file and the VM's is untouched. The server starts with an empty ruleset:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-web nft list ruleset
$ sudo ip netns exec hard-nft-web ss -tln
State Recv-Q Send-Q Local Address:Port Peer Address:Port LISTEN 0 5 0.0.0.0:22 0.0.0.0:* LISTEN 0 5 0.0.0.0:80 0.0.0.0:* LISTEN 0 5 0.0.0.0:9100 0.0.0.0:*

Everything listens on every address. The firewall decides who reaches what. Here is the file:

/etc/nftables.conf
#!/usr/sbin/nft -f
# Host firewall for web01: default-deny input for IPv4 and IPv6.
# Loaded by nftables.service at boot; test changes with nft -c -f first.
# Start from an empty ruleset, so this file is the complete firewall.
flush ruleset
table inet filter {
# Networks allowed to reach SSH (IPv4). Replace 192.0.2.0/24 with your own
# management networks before loading; add an ipv6_addr set for IPv6.
set admin_v4 {
type ipv4_addr
flags interval
elements = { 192.0.2.0/24 }
}
chain input {
type filter hook input priority filter; policy drop;
# The host talking to itself.
iif "lo" accept
# Replies to connections this host opened, and their ICMP errors.
ct state established,related accept
# Packets that belong to no connection conntrack knows.
ct state invalid counter drop
# ICMP and ICMPv6: ping, path MTU discovery, IPv6 neighbour discovery.
meta l4proto { icmp, ipv6-icmp } accept
# SSH from the admin networks only.
tcp dport 22 ip saddr @admin_v4 counter accept
# The public web service.
tcp dport { 80, 443 } counter accept
# Log a sample of what is about to be dropped, and count all of it.
limit rate 5/minute burst 5 packets log prefix "nft-drop: " level info
counter comment "dropped by policy"
}
chain forward {
type filter hook forward priority filter; policy drop;
}
chain output {
type filter hook output priority filter; policy accept;
}
}

The pieces, top to bottom. flush ruleset deletes every table first, so the file describes the whole firewall. A table holds chains and sets; the inet family covers IPv4 and IPv6 in one table. admin_v4 is a named set of IPv4 addresses; flags interval lets it hold ranges. 192.0.2.0/24 is a documentation range standing in for the lab's admin network: before you load this file on a real server, put the addresses you really manage it from in the set, or the first load cuts off every new SSH login. The input chain is attached to the input hook, where packets addressed to this host arrive, at the standard filter priority, and policy drop is what happens to a packet that no rule accepts.

Rules are checked in order and the first verdict wins. Loopback and packets belonging to connections this host already has (ct state established,related, connection tracking) are accepted first; those are almost all packets on a busy host. Each tracked connection is an entry in a fixed-size table: on a host with very many connections (a load balancer, a DNS server), compare net.netfilter.nf_conntrack_count with nf_conntrack_max before you enable a stateful ruleset, because a full table drops new connections and the kernel logs "nf_conntrack: table full, dropping packet" (kernel nf_conntrack-sysctl documentation). meta l4proto { icmp, ipv6-icmp } accepts ICMP for both families; l4proto also finds ICMPv6 behind IPv6 extension headers, which ip6 nexthdr ipv6-icmp would miss (nft(8)), and IPv6 does not work without neighbour discovery. SSH is accepted only from @admin_v4, which also makes SSH IPv4-only: add an ipv6_addr set and an ip6 saddr rule for an IPv6 management network. The web ports are open to all. Whatever reaches the end is logged (at most five packets a minute, so a scan cannot flood the log), counted, and dropped by the policy. forward drops everything because this server routes nothing, and output allows everything this host sends. nft accepts comments at the end of a line too; this file keeps them on their own lines like every configuration in this course.

One inbound packet through the input chain
1Packet addressed to this host
input hook, rules in order
2Loopback or known connection?
accept: iif lo, ct state
3Invalid, or ICMP?
drop invalid, accept ICMP and ICMPv6
4SSH from @admin_v4?
accept, counted
5Web port 80 or 443?
accept, counted
6Nothing matched
log a sample, count, policy drop
First match wins; the policy is the last word.

Check, load, and prove it from outside

nft -c -f parses the file and checks it against the kernel without changing anything; silence means it is valid. nft -f loads it.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-web nft -c -f /etc/nftables.conf
$ sudo ip netns exec hard-nft-web nft -f /etc/nftables.conf sudo ip netns exec hard-nft-web nft list ruleset
table inet filter { set admin_v4 { type ipv4_addr flags interval elements = { 192.0.2.0/24 } } chain input { type filter hook input priority filter; policy drop; iif "lo" accept ct state established,related accept ct state invalid counter packets 0 bytes 0 drop meta l4proto { icmp, ipv6-icmp } accept tcp dport 22 ip saddr @admin_v4 counter packets 0 bytes 0 accept tcp dport { 80, 443 } counter packets 0 bytes 0 accept limit rate 5/minute burst 5 packets log prefix "nft-drop: " level info counter packets 0 bytes 0 comment "dropped by policy" } chain forward { type filter hook forward priority filter; policy drop; } chain output { type filter hook output priority filter; policy accept; } }

nft list ruleset prints what the kernel holds, not what the file says. The counters appear with zeros, and the set shows its range. Now test from both neighbours. nc -z only opens a TCP connection, and -w 3 gives up after three seconds.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-adm nc -zv -w 3 192.0.2.1 22
Connection to 192.0.2.1 22 port [tcp/ssh] succeeded!
$ sudo ip netns exec hard-nft-out nc -zv -w 3 203.0.113.1 80 sudo ip netns exec hard-nft-out nc -zv -w 3 203.0.113.1 22 sudo ip netns exec hard-nft-out nc -zv -w 3 203.0.113.1 9100
Connection to 203.0.113.1 80 port [tcp/http] succeeded! nc: connect to 203.0.113.1 port 22 (tcp) timed out: Operation now in progress nc: connect to 203.0.113.1 port 9100 (tcp) timed out: Operation now in progress
$ sudo ip netns exec hard-nft-web nft list chain inet filter input
table inet filter { chain input { type filter hook input priority filter; policy drop; iif "lo" accept ct state established,related accept ct state invalid counter packets 0 bytes 0 drop meta l4proto { icmp, ipv6-icmp } accept tcp dport 22 ip saddr @admin_v4 counter packets 1 bytes 60 accept tcp dport { 80, 443 } counter packets 1 bytes 60 accept limit rate 5/minute burst 5 packets log prefix "nft-drop: " level info counter packets 6 bytes 360 comment "dropped by policy" } }

The admin host reaches SSH. From outside, the web port answers, and SSH and the metrics port time out: a dropped packet gets no answer at all, so the client waits, unlike "Connection refused", which means a host answered that nothing listens. The counters agree: one SSH packet from the admin host, one web packet, and six packets that reached the policy, because each of the two blocked connection attempts sent its first SYN and two retransmissions before nc gave up. The log shows who knocked:

deploy@web01 · Ubuntu 26.04 LTS
$ journalctl -k --since '-10min' --no-hostname --grep 'nft-drop: .*SRC=203.0.113.50' | tail -n 2
Sep 27 10:04:18 kernel: nft-drop: IN=nft-web-b OUT= MAC=c2:b2:c7:3c:d3:82:16:87:ea:b8:16:3f:08:00 SRC=203.0.113.50 DST=203.0.113.1 LEN=60 TOS=0x00 PREC=0x00 TTL=64 ID=24778 DF PROTO=TCP SPT=59170 DPT=22 WINDOW=64240 RES=0x00 SYN URGP=0 Sep 27 10:04:19 kernel: nft-drop: IN=nft-web-b OUT= MAC=c2:b2:c7:3c:d3:82:16:87:ea:b8:16:3f:08:00 SRC=203.0.113.50 DST=203.0.113.1 LEN=60 TOS=0x00 PREC=0x00 TTL=64 ID=24779 DF PROTO=TCP SPT=59170 DPT=22 WINDOW=64240 RES=0x00 SYN URGP=0

Each line carries the interface (IN=), source and destination addresses, and ports; SYN marks a connection attempt, and the same source port (SPT) on both lines shows one attempt retransmitted. The kernel logs netfilter messages only from the host's own network namespace unless net.netfilter.nf_log_all_netns is 1 (kernel netfilter sysctl documentation), so the lab set it for this demonstration and restored it; on a real server's ruleset you do not need it.

Changing a live ruleset safely

The monitoring host needs port 9100. Add a rule to the file before the web rule, with a typo, and load it without checking first:

/etc/nftables.conf (added before the web service rule)
# Metrics, scraped by the monitoring host only.
tcp dprot 9100 ip saddr 192.0.2.10 counter accept
deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-web nft -f /etc/nftables.conf
/etc/nftables.conf:31:7-11: Error: syntax error, unexpected string tcp dprot 9100 ip saddr 192.0.2.10 counter accept ^^^^^
$ sudo ip netns exec hard-nft-web nft list chain inet filter input | grep -c accept
5

nft points at line 31, column 7. The running ruleset still has its five accept rules, so the flush ruleset on line 6 did not happen either: nft sends a file to the kernel as one transaction, applied completely or not at all (nftables wiki, "Atomic rule replacement"). A firewall is never left half-loaded. With the typo fixed, the check is silent, the load succeeds, and the monitoring host gets through:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-web nft -c -f /etc/nftables.conf sudo ip netns exec hard-nft-web nft -f /etc/nftables.conf sudo ip netns exec hard-nft-adm nc -zv -w 3 192.0.2.1 9100
Connection to 192.0.2.1 9100 port [tcp/*] succeeded!

Sets can change without any reload. Adding an address to admin_v4 lets the outside host in at once, and deleting it closes the door again:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-web nft add element inet filter admin_v4 '{ 203.0.113.50 }' sudo ip netns exec hard-nft-web nft list set inet filter admin_v4 sudo ip netns exec hard-nft-out nc -zv -w 3 203.0.113.1 22
table inet filter { set admin_v4 { type ipv4_addr flags interval elements = { 192.0.2.0/24, 203.0.113.50 } } } Connection to 203.0.113.1 22 port [tcp/ssh] succeeded!
$ sudo ip netns exec hard-nft-web nft delete element inet filter admin_v4 '{ 203.0.113.50 }' sudo ip netns exec hard-nft-out nc -zv -w 3 203.0.113.1 22
nc: connect to 203.0.113.1 port 22 (tcp) timed out: Operation now in progress

That is how you grant temporary access or block an attacker without touching the rules. A change made with nft add lives only in the kernel, though: the next load of the file or reboot removes it. Put anything permanent in the file.

A rollback that actually runs

Changing the firewall of a remote server over SSH is the classic way to lock yourself out. The established-connection rule usually keeps your current session alive, but the next login fails. So before a risky change, arm something that restores the known-good ruleset on its own unless you cancel it. An old recipe does this with echo 'flush ruleset' | at now + 5 minutes. It fails twice: at runs the text with /bin/sh (at(1)), where flush is not a command, and at is not installed on Ubuntu 26.04 at all:

deploy@web01 · Ubuntu 26.04 LTS
$ command -v at

systemd-run --on-active= is always available. It creates a transient timer that starts a transient service after the given time. Here it reloads /etc/nftables.conf, the known-good file, two minutes from now; on a real server it is sudo systemd-run --on-active=5min --unit=nft-rollback nft -f /etc/nftables.conf, and five minutes or more gives you time to test. The very first load has no known-good file to go back to; arm sudo systemd-run --on-active=5min --unit=nft-rollback nft flush ruleset instead, which returns the host to no firewall at all (on a host where other tools own tables, nft destroy table inet filter removes only yours).

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --on-active=2min --unit=hard-nft-rollback ip netns exec hard-nft-web nft -f /etc/nftables.conf
Running timer as unit: hard-nft-rollback.timer Will run service as unit: hard-nft-rollback.service
$ systemctl list-timers hard-nft-rollback.timer
NEXT LEFT LAST PASSED UNIT ACTIVATES Sun 2026-09-27 10:06:21 UTC 1min 59s - - hard-nft-rollback.timer hard-nft-rollback.service 1 timers listed. Pass --all to see loaded but inactive timers, too.

Now the risky change: replace the SSH allow-list with a new management range, typed wrongly. The admin host is locked out at once:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec hard-nft-web nft 'flush set inet filter admin_v4; add element inet filter admin_v4 { 198.51.100.0/24 }'
$ sudo ip netns exec hard-nft-adm nc -zv -w 3 192.0.2.1 22
nc: connect to 192.0.2.1 port 22 (tcp) timed out: Operation now in progress

Two minutes later the timer fires, the service reloads the file, and the admin host is back in:

deploy@web01 · Ubuntu 26.04 LTS
$ journalctl -u hard-nft-rollback --since -4min --no-hostname sudo ip netns exec hard-nft-web nft list set inet filter admin_v4 sudo ip netns exec hard-nft-adm nc -zv -w 3 192.0.2.1 22
Sep 27 10:06:21 systemd[1]: Started hard-nft-rollback.service - [systemd-run] /usr/sbin/ip netns exec hard-nft-web nft -f /etc/nftables.conf. Sep 27 10:06:21 systemd[1]: hard-nft-rollback.service: Deactivated successfully. table inet filter { set admin_v4 { type ipv4_addr flags interval elements = { 192.0.2.0/24 } } } Connection to 192.0.2.1 22 port [tcp/ssh] succeeded!

When a change works, cancel the rollback by stopping the timer; list-timers then shows nothing left to fire:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --on-active=5min --unit=hard-nft-rollback ip netns exec hard-nft-web nft -f /etc/nftables.conf sudo systemctl stop hard-nft-rollback.timer systemctl list-timers hard-nft-rollback.timer
Running timer as unit: hard-nft-rollback.timer Will run service as unit: hard-nft-rollback.service NEXT LEFT LAST PASSED UNIT ACTIVATES 0 timers listed. Pass --all to see loaded but inactive timers, too.

Change the live ruleset first and the file only after the change has proved itself. If you do edit the file first, copy the known-good version aside (sudo cp /etc/nftables.conf /etc/nftables.conf.good) and point the rollback at the copy. The rollback is SecOpsLog advice; its impact is only the minutes you have to confirm in, so keep a second SSH session open while you test.

Loading it at boot, and one owner per host

nftables.service loads the file at boot. Both platforms ship it disabled, and they read different files:

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl cat nftables.service | grep -E '^(Exec|Wanted)' systemctl is-enabled nftables.service
ExecStart=/usr/sbin/nft -f /etc/nftables.conf ExecReload=/usr/sbin/nft -f /etc/nftables.conf ExecStop=/usr/sbin/nft flush ruleset WantedBy=sysinit.target disabled
deploy@rocky10 · Rocky Linux 10.2
$ systemctl cat nftables.service | grep -E '^Exec' systemctl is-enabled nftables.service
ExecStart=/sbin/nft -f /etc/sysconfig/nftables.conf ExecReload=/sbin/nft 'flush ruleset; include "/etc/sysconfig/nftables.conf";' ExecStop=/sbin/nft flush ruleset disabled
$ sudo grep -v "^$" /etc/sysconfig/nftables.conf
# Uncomment the include statement here to load the default config sample # in /etc/nftables for nftables service. #include "/etc/nftables/main.nft" # To customize, either edit the samples in /etc/nftables, append further # commands to the end of this file or overwrite it after first service # start by calling: 'nft list ruleset >/etc/sysconfig/nftables.conf'.

Ubuntu loads /etc/nftables.conf; its shipped file is an empty inet filter table with no policy, which accepts everything. RHEL loads /etc/sysconfig/nftables.conf, which ships as comments only, and RHEL's reload runs flush ruleset itself before including the file. Ubuntu's reload runs the file as it is, so keep flush ruleset at its top, or every reload adds a second copy of each rule. WantedBy=sysinit.target with Before=network-pre.target (in the unit) loads the rules before the network is up. To use your file, put it in place, check it with sudo nft -c -f, and run sudo systemctl enable --now nftables.service. The rollback is sudo systemctl disable --now nftables.service.

A ruleset that fails to load at boot fails open: nftables.service is marked failed, the kernel has no ruleset, and every port is reachable. Two habits prevent it. Never put host names in the file (resolving them needs DNS, and the unit runs before the network is up), and check every edit with nft -c -f before it is saved for the next boot. Then alert on systemctl is-failed nftables.service from your monitoring, the same way you would for any unit that must be running.

Both units stop with nft flush ruleset, which deletes every table on the host: ufw's, firewalld's, Docker's, and on these lab machines the VM tool's own NAT table. A file that starts with flush ruleset does the same at every reload. On a container or virtualisation host, where the runtime owns tables of its own, start the file with destroy table inet filter instead: nft(8) documents destroy as a delete that does not fail when the table does not exist (nftables 1.0.7 and later, with a Linux 6.3 or newer kernel; the lab has nftables 1.1.6), so each load replaces only your table, still in one transaction, and the runtime's tables survive; the lab checked that another table was left in place. That is also why a raw ruleset and a front-end never share a host, and systemd will not stop you enabling both (see "Host firewalls: ufw and firewalld"). Choose one owner per host: firewalld on RHEL or ufw on Ubuntu when their zones and services cover your needs, a hand-written nftables ruleset when you need sets, counters or precise control, and disable the others.

Try this

On your Ubuntu lab machine, install socat, save the script above and run it with sudo. Put this lesson's ruleset in /etc/netns/hard-nft-web/nftables.conf and load it with sudo ip netns exec hard-nft-web nft -f /etc/nftables.conf. From the two neighbours, confirm with nc -zv -w 3 that SSH answers only the admin host. Arm a two-minute rollback with sudo systemd-run --on-active=2min ip netns exec hard-nft-web nft -f /etc/nftables.conf, empty admin_v4 with nft flush set inet filter admin_v4 (inside the namespace), watch the admin host's test connection time out, and watch it come back when the timer fires. Clean up with sudo systemctl stop hard-nft-l22 hard-nft-l80 hard-nft-l9100, sudo ip netns del for each of the three namespaces, and sudo rm -r /etc/netns/hard-nft-web.

Takeaway

Keep the whole firewall in one file that replaces the old ruleset in one step (flush ruleset, or destroy table of your own table where another tool owns tables), check it with nft -c, load it with nft -f (all or nothing), and prove it from another host. Before any remote change, arm a systemd-run --on-active rollback to the known-good file, and let exactly one tool own the host's ruleset.

Quick check
01Your ruleset accepts tcp dport 22 ip saddr @admin_v4 in an inet table. An engineer on the IPv6 management network 2001:db8:10::/64 cannot reach SSH. Why?
Incorrect — The inet family covers both IPv4 and IPv6; the ICMPv6 rule in the same table works for IPv6.
Incorrect — Conntrack tracks IPv6 as well; established,related accepts replies for both families.
Correct — An ip saddr match implies IPv4. Add an ipv6_addr set with the management prefix and a matching ip6 saddr rule.
Incorrect — The policy only applies after every rule in the chain has been checked, for both families.
02Before tightening the SSH allow-list on a remote server you run: sudo systemd-run --on-active=5min --unit=nft-rollback nft -f /etc/nftables.conf. The new list works. What is left to do?
Correct — Stopping the timer cancels the rollback; writing the change into the file makes it survive the next load and reboot.
Incorrect — It would reload the old file in five minutes and remove your new list, which lives only in the kernel.
Incorrect — Transient units live only in memory; daemon-reload does not cancel a running timer.
Incorrect — A transient timer is a unit like any other and systemctl stop cancels it.
03nft -f fails on a file whose first command is flush ruleset and whose typo is on line 31. What does the server's firewall look like now?
Incorrect — nft checks the whole file before sending anything, and the kernel applies it as one transaction; no part of it ran.
Correct — The lab's chain still had its five accept rules afterwards. A file is applied completely or not at all.
Incorrect — That is the failure mode of a script of separate commands; one nft -f file is one transaction.
Incorrect — nft has no fallback; the previously loaded ruleset, with its drop policy, stays in place.

Related