Text processing: answering questions from logs
sort, uniq, cut, sed and awk on real logs.
Logs answer questions only when you can cut them down to the part that matters. This lesson works through a web server's access log with a handful of small tools joined by pipes: cut and awk pick out columns, sort and uniq -c count, tr fixes spacing, and sed prints and rewrites lines. By the end you can find which clients send the most requests, which requests failed, whether a client that kept failing ever succeeded, and who moved the most data, and you can prepare a log to share without exposing addresses.
The log and its fields
The examples use ~/text/access.log, a sample log in the format web servers such as nginx and Apache write by default. Its 81 requests come from four clients with addresses from the ranges reserved for documentation, and every time in it is fixed, so your counts will match the ones shown here. Create it by pasting the whole command below into your shell as one block. You do not need to follow it yet: it is an awk program that prints the log lines, and by the end of this lesson you will recognise most of it.
It prints nothing, because its output went into the file. Look at the file's size and first lines before anything else:
Each line is one request. Split at spaces, the fields are: 1 the client address; 2 and 3 identity fields that are usually -; 4 and 5 the time in brackets; 6, 7 and 8 the request (method with a leading quote, path, protocol with a closing quote); 9 the HTTP status code the server answered with; 10 the number of bytes sent back; then the page the client came from and the client's self-description (the user agent), both in quotes. Status codes in the 200s mean success, 302 is a redirect, 401 means the login was refused and 404 that the path does not exist.
Counting one column: cut, sort and uniq -c
cut -d DELIM -f N prints field N of every line, splitting at the delimiter character. The addresses are field 1, separated by spaces:
To count how often each address appears, uniq -c collapses repeated lines into one and puts the count in front. But it only compares each line with the one just before it:
The monitor and the browser take turns in the log, so every short run is counted separately. sort first puts identical lines next to each other; then uniq -c counts them, and a second sort -rn (numeric, reversed) puts the largest count first:
198.51.100.23 sent half of all requests, more than the monitor that checks the site every three minutes. The same pattern works on any file with a fixed delimiter. /etc/passwd separates its fields with colons, and field 7 is the program started when the account logs in. Which login programs does this server use, and how often?
Most accounts belong to services and have nologin (or /bin/false), so nobody can get a shell through them. The few with /bin/bash are root and people; grep bash /etc/passwd names them.
On Ubuntu 26.04 sort, uniq, cut and tr are the Rust uutils versions; for every pipeline in this lesson they printed exactly what the GNU tools (gnusort and the others) print. awk is GNU awk 5.3 and sed is GNU sed, as on RHEL.
Columns padded with spaces: tr -s and awk
Many commands line their output up in columns by padding with several spaces. cut splits at every single space, so the padding turns into empty fields:
Field 5 of both lines is one of the empty strings between padding spaces, so cut prints nothing but empty lines. tr translates or deletes characters, and -s squeezes each run of a repeated character into one, after which cut works. awk needs no help: it splits fields at any run of spaces or tabs.
Two other uses of tr come up often: tr -d '\r' deletes the carriage returns that Windows editors put at the end of every line (the ^M from the lesson on reading files), and tr '[:lower:]' '[:upper:]' changes lower case to upper case.
awk: fields, conditions and sums
An awk program is a list of condition {action} pairs run on every line. $1, $2 and so on are the line's fields and $0 is the whole line. With no condition the action runs on every line; with no action the line is printed. Which status codes did the server send?
Forty refused logins stand out. It is tempting to count them with grep, but grep matches text anywhere on the line, not a field:
grep counted 41 because one successful page was 4012 bytes long. The awk condition $9 == 401 tests the status field only. Now the question that matters: who was refused?
All 40 refusals came from one address, one request every five seconds for over three minutes, which is a program trying passwords rather than a person who forgot one. The next question is whether it ever got in. && combines conditions, != means "not equal", and quoted text is compared as a string:
After forty refusals, one login was answered with 302. A login form usually answers a correct password by redirecting the browser to the next page, so on most sites that line means the password was accepted. Looking for a 200 would have missed it. The next steps are outside this log: find which account logged in (the application's own log), and treat that account's password as known.
awk also keeps variables between lines. bytes[$1] += $10 adds each request's size to a total kept per address in an array, and the END block runs once after the last line to print the totals:
203.0.113.7 sent only 14 requests but received almost 38 million bytes, nearly all of it two downloads of the manual. A client with few requests and many bytes deserves a look, and here the paths show it is harmless. The fourth client is easy to read once you print the address and path of every 404:
A WordPress login page, a .env file with secrets, admin panels and a Git configuration, one per second: an automated scanner checking for common mistakes. None of them exist here, which is the answer you want.
sed: print, replace and edit in place
sed applies editing commands to every line of its input. -n stops it printing lines by default, and /PATTERN/p prints the lines that match, which makes it a grep that can do more:
The most common sed command is substitution, s/OLD/NEW/, which replaces the first match on each line (add g after the last slash for every match). With -E the pattern uses the same extended syntax as grep -E. Before you paste a log into a ticket or a chat, replace what does not need to leave the building, here the client addresses:
^[0-9.]+ means "digits and dots at the start of the line, then a space". It handles IPv4 addresses in the first field only; IPv6 clients and addresses elsewhere on the line (a forwarded-for header, a referring page) pass through untouched, so read the result before you share it. The file itself is unchanged, since sed wrote the result to stdout. -i edits the file in place, and -i.bak keeps the original under a new name. On a copy of the SSH server's configuration:
The inode numbers (the lesson on files and links introduced them) show what -i really does. sed wrote a new file and renamed it over the old name, and the original inode now belongs to the .bak copy. That is usually what you want, but it has consequences: other hard links to the file keep the old contents, and when the name is a symbolic link, GNU sed -i replaces the link with a regular file and leaves its target unchanged, unless you add --follow-symlinks.
sed command without -i first and read the output, or diff it against the file. On a configuration file, keep -i.bak so the previous version is one mv away, and run the service's own check (sshd -t, visudo -c) before you reload it.Try this
Build the counts for two more questions on the sample log. How many requests used each HTTP method? Field 6 holds the method with a leading quote, so print it with awk '{print $6}', strip the quote with tr -d '"', then sort and count (40 GET and 41 POST here). How fast was the failing client going? Keep its lines with grep 198.51.100.23, cut out the hour and minute with cut -d: -f2,3 (the colon splits the timestamp), and count with uniq -c: 12 attempts in each full minute.
Takeaway
Turn every question into "which field, filtered how, counted by what", then build the pipeline one command at a time and check each step's output before adding the next. Use awk when fields are padded or a condition must test one field, and sort before uniq -c, every time.