Text processing: answering questions from logs

sort, uniq, cut, sed and awk on real logs.

Beginner16 min · lesson 8 of 29

Logs answer questions only when you can cut them down to the part that matters. This lesson works through a web server's access log with a handful of small tools joined by pipes: cut and awk pick out columns, sort and uniq -c count, tr fixes spacing, and sed prints and rewrites lines. By the end you can find which clients send the most requests, which requests failed, whether a client that kept failing ever succeeded, and who moved the most data, and you can prepare a log to share without exposing addresses.

The log and its fields

The examples use ~/text/access.log, a sample log in the format web servers such as nginx and Apache write by default. Its 81 requests come from four clients with addresses from the ranges reserved for documentation, and every time in it is fixed, so your counts will match the ones shown here. Create it by pasting the whole command below into your shell as one block. You do not need to follow it yet: it is an awk program that prints the log lines, and by the end of this lesson you will recognise most of it.

deploy@web01 · Ubuntu 26.04 LTS
$ mkdir -p ~/text awk 'function line(s, ip, req, st, b, ua) { printf "%05d %s - - [27/Sep/2026:%02d:%02d:%02d +0000] \"%s HTTP/1.1\" %d %d \"-\" \"%s\"\n", s, ip, 9 + int(s / 3600), int(s % 3600 / 60), s % 60, req, st, b, ua } BEGIN { for (i = 0; i < 20; i++) line(i * 180, "192.0.2.10", "GET /health", 200, 2, "healthcheck/1.0") split("/ /about.html /products.html /products/widget.html /products/gadget.html /pricing.html /contact.html /blog/ /blog/release-notes.html /support.html /faq.html /docs/", pg, " ") split("5120 3301 7640 4012 4388 2950 2204 8811 6127 3478 5902 7019", by, " ") for (i = 1; i <= 12; i++) line(300 + (i - 1) * 97, "203.0.113.7", "GET " pg[i], 200, by[i], "Mozilla/5.0 (X11; Linux x86_64)") line(1500, "203.0.113.7", "GET /downloads/manual.pdf", 200, 18874368, "Mozilla/5.0 (X11; Linux x86_64)") line(1510, "203.0.113.7", "GET /downloads/manual.pdf", 200, 18874368, "Mozilla/5.0 (X11; Linux x86_64)") split("/wp-login.php /.env /admin/ /phpmyadmin/ /.git/config", pr, " ") for (i = 1; i <= 5; i++) line(2000 + i, "192.0.2.44", "GET " pr[i], 404, 162, "python-requests/2.32") line(2006, "192.0.2.44", "GET /login", 200, 1834, "python-requests/2.32") for (i = 0; i < 40; i++) line(2400 + i * 5, "198.51.100.23", "POST /login", 401, 512, "Go-http-client/1.1") line(2605, "198.51.100.23", "POST /login", 302, 0, "Go-http-client/1.1") }' | sort -n | cut -d" " -f2- > ~/text/access.log

It prints nothing, because its output went into the file. Look at the file's size and first lines before anything else:

deploy@web01 · Ubuntu 26.04 LTS
$ wc -l ~/text/access.log
81 /home/deploy/text/access.log
$ head -n 3 ~/text/access.log
192.0.2.10 - - [27/Sep/2026:09:00:00 +0000] "GET /health HTTP/1.1" 200 2 "-" "healthcheck/1.0" 192.0.2.10 - - [27/Sep/2026:09:03:00 +0000] "GET /health HTTP/1.1" 200 2 "-" "healthcheck/1.0" 203.0.113.7 - - [27/Sep/2026:09:05:00 +0000] "GET / HTTP/1.1" 200 5120 "-" "Mozilla/5.0 (X11; Linux x86_64)"

Each line is one request. Split at spaces, the fields are: 1 the client address; 2 and 3 identity fields that are usually -; 4 and 5 the time in brackets; 6, 7 and 8 the request (method with a leading quote, path, protocol with a closing quote); 9 the HTTP status code the server answered with; 10 the number of bytes sent back; then the page the client came from and the client's self-description (the user agent), both in quotes. Status codes in the 200s mean success, 302 is a redirect, 401 means the login was refused and 404 that the path does not exist.

Counting one column: cut, sort and uniq -c

cut -d DELIM -f N prints field N of every line, splitting at the delimiter character. The addresses are field 1, separated by spaces:

deploy@web01 · Ubuntu 26.04 LTS
$ cut -d" " -f1 ~/text/access.log | head -n 4
192.0.2.10 192.0.2.10 203.0.113.7 192.0.2.10

To count how often each address appears, uniq -c collapses repeated lines into one and puts the count in front. But it only compares each line with the one just before it:

deploy@web01 · Ubuntu 26.04 LTS
$ cut -d" " -f1 ~/text/access.log | uniq -c | head -n 6
2 192.0.2.10 1 203.0.113.7 1 192.0.2.10 2 203.0.113.7 1 192.0.2.10 2 203.0.113.7

The monitor and the browser take turns in the log, so every short run is counted separately. sort first puts identical lines next to each other; then uniq -c counts them, and a second sort -rn (numeric, reversed) puts the largest count first:

deploy@web01 · Ubuntu 26.04 LTS
$ cut -d" " -f1 ~/text/access.log | sort | uniq -c | sort -rn
41 198.51.100.23 20 192.0.2.10 14 203.0.113.7 6 192.0.2.44
Counting requests per client
1cut -d" " -f1
keep the address column
2sort
bring identical addresses together
3uniq -c
one line per address, with a count
4sort -rn
largest count first
Each command reads the previous one's output. Add head -n 5 at the end when the list is long.

198.51.100.23 sent half of all requests, more than the monitor that checks the site every three minutes. The same pattern works on any file with a fixed delimiter. /etc/passwd separates its fields with colons, and field 7 is the program started when the account logs in. Which login programs does this server use, and how often?

deploy@web01 · Ubuntu 26.04 LTS
$ cut -d: -f7 /etc/passwd | sort | uniq -c | sort -rn
30 /usr/sbin/nologin 3 /bin/bash 1 /bin/sync 1 /bin/false

Most accounts belong to services and have nologin (or /bin/false), so nobody can get a shell through them. The few with /bin/bash are root and people; grep bash /etc/passwd names them.

On Ubuntu 26.04 sort, uniq, cut and tr are the Rust uutils versions; for every pipeline in this lesson they printed exactly what the GNU tools (gnusort and the others) print. awk is GNU awk 5.3 and sed is GNU sed, as on RHEL.

Columns padded with spaces: tr -s and awk

Many commands line their output up in columns by padding with several spaces. cut splits at every single space, so the padding turns into empty fields:

deploy@web01 · Ubuntu 26.04 LTS
$ df -h /
Filesystem Size Used Avail Use% Mounted on /dev/vda1 23G 4.2G 19G 19% /
$ df -h / | cut -d" " -f5

Field 5 of both lines is one of the empty strings between padding spaces, so cut prints nothing but empty lines. tr translates or deletes characters, and -s squeezes each run of a repeated character into one, after which cut works. awk needs no help: it splits fields at any run of spaces or tabs.

deploy@web01 · Ubuntu 26.04 LTS
$ df -h / | tr -s " " | cut -d" " -f5
Use% 19%
$ df -h / | awk '{print $5}'
Use% 19%

Two other uses of tr come up often: tr -d '\r' deletes the carriage returns that Windows editors put at the end of every line (the ^M from the lesson on reading files), and tr '[:lower:]' '[:upper:]' changes lower case to upper case.

awk: fields, conditions and sums

An awk program is a list of condition {action} pairs run on every line. $1, $2 and so on are the line's fields and $0 is the whole line. With no condition the action runs on every line; with no action the line is printed. Which status codes did the server send?

deploy@web01 · Ubuntu 26.04 LTS
$ awk '{print $9}' ~/text/access.log | sort | uniq -c
35 200 1 302 40 401 5 404

Forty refused logins stand out. It is tempting to count them with grep, but grep matches text anywhere on the line, not a field:

deploy@web01 · Ubuntu 26.04 LTS
$ grep -c 401 ~/text/access.log
41
$ awk '$9 == 401' ~/text/access.log | wc -l
40
$ grep 401 ~/text/access.log | grep -v '" 401 '
203.0.113.7 - - [27/Sep/2026:09:09:51 +0000] "GET /products/widget.html HTTP/1.1" 200 4012 "-" "Mozilla/5.0 (X11; Linux x86_64)"

grep counted 41 because one successful page was 4012 bytes long. The awk condition $9 == 401 tests the status field only. Now the question that matters: who was refused?

deploy@web01 · Ubuntu 26.04 LTS
$ awk '$9 == 401 {print $1}' ~/text/access.log | sort | uniq -c | sort -rn
40 198.51.100.23

All 40 refusals came from one address, one request every five seconds for over three minutes, which is a program trying passwords rather than a person who forgot one. The next question is whether it ever got in. && combines conditions, != means "not equal", and quoted text is compared as a string:

deploy@web01 · Ubuntu 26.04 LTS
$ awk '$1 == "198.51.100.23" && $9 != 401' ~/text/access.log
198.51.100.23 - - [27/Sep/2026:09:43:25 +0000] "POST /login HTTP/1.1" 302 0 "-" "Go-http-client/1.1"

After forty refusals, one login was answered with 302. A login form usually answers a correct password by redirecting the browser to the next page, so on most sites that line means the password was accepted. Looking for a 200 would have missed it. The next steps are outside this log: find which account logged in (the application's own log), and treat that account's password as known.

awk also keeps variables between lines. bytes[$1] += $10 adds each request's size to a total kept per address in an array, and the END block runs once after the last line to print the totals:

deploy@web01 · Ubuntu 26.04 LTS
$ awk '{bytes[$1] += $10} END {for (ip in bytes) print bytes[ip], ip}' ~/text/access.log | sort -rn
37809688 203.0.113.7 20480 198.51.100.23 2644 192.0.2.44 40 192.0.2.10

203.0.113.7 sent only 14 requests but received almost 38 million bytes, nearly all of it two downloads of the manual. A client with few requests and many bytes deserves a look, and here the paths show it is harmless. The fourth client is easy to read once you print the address and path of every 404:

deploy@web01 · Ubuntu 26.04 LTS
$ awk '$9 == 404 {print $1, $7}' ~/text/access.log
192.0.2.44 /wp-login.php 192.0.2.44 /.env 192.0.2.44 /admin/ 192.0.2.44 /phpmyadmin/ 192.0.2.44 /.git/config

A WordPress login page, a .env file with secrets, admin panels and a Git configuration, one per second: an automated scanner checking for common mistakes. None of them exist here, which is the answer you want.

sed: print, replace and edit in place

sed applies editing commands to every line of its input. -n stops it printing lines by default, and /PATTERN/p prints the lines that match, which makes it a grep that can do more:

deploy@web01 · Ubuntu 26.04 LTS
$ sed -n '/192.0.2.44/p' ~/text/access.log
192.0.2.44 - - [27/Sep/2026:09:33:21 +0000] "GET /wp-login.php HTTP/1.1" 404 162 "-" "python-requests/2.32" 192.0.2.44 - - [27/Sep/2026:09:33:22 +0000] "GET /.env HTTP/1.1" 404 162 "-" "python-requests/2.32" 192.0.2.44 - - [27/Sep/2026:09:33:23 +0000] "GET /admin/ HTTP/1.1" 404 162 "-" "python-requests/2.32" 192.0.2.44 - - [27/Sep/2026:09:33:24 +0000] "GET /phpmyadmin/ HTTP/1.1" 404 162 "-" "python-requests/2.32" 192.0.2.44 - - [27/Sep/2026:09:33:25 +0000] "GET /.git/config HTTP/1.1" 404 162 "-" "python-requests/2.32" 192.0.2.44 - - [27/Sep/2026:09:33:26 +0000] "GET /login HTTP/1.1" 200 1834 "-" "python-requests/2.32"

The most common sed command is substitution, s/OLD/NEW/, which replaces the first match on each line (add g after the last slash for every match). With -E the pattern uses the same extended syntax as grep -E. Before you paste a log into a ticket or a chat, replace what does not need to leave the building, here the client addresses:

deploy@web01 · Ubuntu 26.04 LTS
$ sed -E 's/^[0-9.]+ /203.0.113.x /' ~/text/access.log | head -n 2
203.0.113.x - - [27/Sep/2026:09:00:00 +0000] "GET /health HTTP/1.1" 200 2 "-" "healthcheck/1.0" 203.0.113.x - - [27/Sep/2026:09:03:00 +0000] "GET /health HTTP/1.1" 200 2 "-" "healthcheck/1.0"

^[0-9.]+ means "digits and dots at the start of the line, then a space". It handles IPv4 addresses in the first field only; IPv6 clients and addresses elsewhere on the line (a forwarded-for header, a referring page) pass through untouched, so read the result before you share it. The file itself is unchanged, since sed wrote the result to stdout. -i edits the file in place, and -i.bak keeps the original under a new name. On a copy of the SSH server's configuration:

deploy@web01 · Ubuntu 26.04 LTS
$ cp /etc/ssh/sshd_config ~/text/sshd_config grep -n MaxAuthTries ~/text/sshd_config ls -i ~/text/sshd_config
56:#MaxAuthTries 6 575527 /home/deploy/text/sshd_config
$ sed -i.bak 's/^#MaxAuthTries 6/MaxAuthTries 3/' ~/text/sshd_config grep -n MaxAuthTries ~/text/sshd_config ls -i ~/text/sshd_config ~/text/sshd_config.bak
56:MaxAuthTries 3 575528 /home/deploy/text/sshd_config 575527 /home/deploy/text/sshd_config.bak

The inode numbers (the lesson on files and links introduced them) show what -i really does. sed wrote a new file and renamed it over the old name, and the original inode now belongs to the .bak copy. That is usually what you want, but it has consequences: other hard links to the file keep the old contents, and when the name is a symbolic link, GNU sed -i replaces the link with a regular file and leaves its target unchanged, unless you add --follow-symlinks.

Preview before -i
Run the sed command without -i first and read the output, or diff it against the file. On a configuration file, keep -i.bak so the previous version is one mv away, and run the service's own check (sshd -t, visudo -c) before you reload it.

Try this

Build the counts for two more questions on the sample log. How many requests used each HTTP method? Field 6 holds the method with a leading quote, so print it with awk '{print $6}', strip the quote with tr -d '"', then sort and count (40 GET and 41 POST here). How fast was the failing client going? Keep its lines with grep 198.51.100.23, cut out the hour and minute with cut -d: -f2,3 (the colon splits the timestamp), and count with uniq -c: 12 attempts in each full minute.

Takeaway

Turn every question into "which field, filtered how, counted by what", then build the pipeline one command at a time and check each step's output before adding the next. Use awk when fields are padded or a condition must test one field, and sort before uniq -c, every time.

Quick check
01You run cut -d" " -f1 access.log | uniq -c on a busy log, and the same address appears on many lines, each with a small count. What went wrong?
Incorrect — uniq compares whole lines as text. Addresses are counted like anything else.
Incorrect — cut prints the field without its delimiter. The lines really are identical; they are just not adjacent.
Incorrect — uniq streams its input line by line and has no such limit. The splits follow the order of the lines.
Correct — uniq compares each line with the previous one; sort puts every copy of an address together.
02In a script, df -h / | cut -d" " -f5 prints blank lines instead of the Use% column. What fixes it?
Incorrect — df pads its columns with spaces. Splitting at tabs finds no delimiter and prints whole lines.
Correct — cut splits at every single space, so padding produces empty fields; tr -s or awk treats a run of spaces as one separator.
Incorrect — df shows the same figures to every user. The fields are there; cut is splitting them wrongly.
Incorrect — 5- means field 5 to the end of the line. Field 5 is still an empty string between padding spaces.
03/etc/app.conf is a symbolic link to /srv/app/app.conf. You run sudo sed -i "s/info/debug/" /etc/app.conf, and the application, which reads /srv/app/app.conf, still logs at info. Why?
Correct — sed -i writes a new file and renames it over the name; for a link that replaces the link itself unless --follow-symlinks is given.
Incorrect — sed changes files immediately. daemon-reload concerns systemd unit files, and the application reads a different file anyway.
Incorrect — sed ran as root and wrote a file; ls -l /etc/app.conf would show it is no longer a link.
Incorrect — Without g, sed replaces the first match on each line, which is enough for level=info.

Related