Archives, downloads and checksums
tar, compression, rsync and verifying files.
Sooner or later you need to move files: a copy of a configuration directory before you change it, a week of logs for a colleague, a release onto a server, a backup onto another machine. This lesson covers the tools for each part of that job: tar to bundle a directory into one file, gzip, xz and zstd to make it smaller, rsync to copy it to another machine over SSH, and curl with sha256sum and gpgv to download a file and prove it is the one its publisher released. The last part matters most for security: a download you have not verified is code you run on trust.
Bundling a directory with tar
tar packs many files and directories into one archive file, keeping their names, permissions, owners and modification times, and unpacks them again. It does not compress anything by itself; an option passes the archive through a compressor on the way. The practice directory for this lesson is a small web application, ~/cmd-archive/webapp, with a configuration file, a page and two logs. The commands below create it; the awk line writes 120,000 made-up access-log lines, and its details do not matter here:
The terminals in this lesson start each step with cd ~/cmd-archive, which is harmless to repeat in your own terminal.
du -ah (disk usage, all files, human-readable sizes) shows that nearly all of it is access.log. tar's options are single letters: -c creates an archive, -t lists one, -x extracts one, -f NAME names the archive file, -v prints each file as it goes, and -z compresses with gzip. By convention a gzip-compressed tar archive ends in .tar.gz.
Creating the archive prints nothing, which means it worked. The listing reads like ls -l: permissions, owner and group, size in bytes, modification time and name. Every name starts with webapp/ because that is how the directory was named on the command line, and each directory comes before the files inside it. List any archive you did not make before you extract it, so you know where its files will land.
-C DIR extracts into another directory instead of the current one. Extracting needs no -z: GNU tar recognises the compression format when it reads an archive.
When you archive files by their absolute path, tar removes the leading slash and says so:
The members are stored as etc/ssh/..., so extracting them can never overwrite the real /etc: they land under the directory you extract into, and you copy back what you need. The second message applies the same rule to hard links. (10-acceptenv-colorterm.conf is added by the lab's VM tool, not by Ubuntu.) The listing also shows that the archive recorded root/root as the owner. What happens to that depends on who extracts it:
Extracted by deploy, the file belongs to deploy: an ordinary user always extracts files as themselves. Run as root (sudo tar -x), tar restores the owners and permissions recorded in the archive, which tar(1) lists as the default for the superuser. That is what you want when you restore your own backup of /etc, and dangerous with an archive from anyone else, because it can create root-owned or set-user-ID files wherever it is unpacked. List first, extract as yourself into an empty directory, and use sudo only for archives you made.
Compressing: gzip, xz and zstd
Compression finds repeated patterns and stores them once, and logs repeat themselves a lot. gzip FILE replaces the file with a compressed FILE.gz; -k keeps the original, and gunzip (or gzip -d) reverses it.
The 7.8 MB log became 1.2 MB. Ubuntu Server 26.04 also installs xz and zstd (Zstandard), which tar uses with -J and --zstd. time -p in front of a command reports how long it took; real is the elapsed time in seconds.
Even this small example shows the trade-off. xz made the smallest archive, a third smaller than gzip's, but took about 2.4 seconds against less than a tenth of a second. zstd at its default level was the fastest and within 5% of gzip's size. A reasonable rule: gzip when the archive has to open anywhere, zstd for large backups where time matters, and xz for files compressed once and downloaded many times, such as release archives. All three are lossless: the files come back byte for byte.
A file name's ending is only a convention. file reads the first bytes and names the real format, and tar needs no hint to read any of them:
Not every archiver is installed by default. On a fresh Ubuntu Server 26.04:
command -v prints the path of each command it finds and nothing for the others, so zip and unzip are missing. You need them only to exchange files with people on Windows or macOS, where .zip opens with a double-click: sudo apt install zip unzip, then zip -r site.zip site. Do not rely on zip's password option (-e) for secrecy; zip(1) itself calls the standard zip encryption relatively weak, so encrypt with a proper tool such as gpg instead. On RHEL 10, tar, gzip and xz are present, but the Rocky Linux 10.2 cloud image used for this course's cross-checks has no zstd command until you install the zstd package.
Copying to another machine with rsync
rsync copies files and directories, locally or to another machine over SSH, and compares both sides first so that it sends only what changed. It has to be installed on both machines, and Ubuntu Server includes it. SSH handles the login and the encryption, so everything from "Connecting with SSH" applies: keys, the host key prompt on the first connection, and host aliases in ~/.ssh/config. The lab has only one machine, so the backup server is a second account on it, ca-backup, reached over SSH at 127.0.0.1. Create the account, a key without a passphrase for it (a backup job has nobody to type one), and install the public key in the account's authorized_keys with the right owner and modes (install -d creates a directory, install copies a file, both setting -m mode and -o/-g owner and group):
Then add a host alias backup01 to ~/.ssh/config with an editor. The first connection asks about the host key, as in "Connecting with SSH"; for 127.0.0.1 it is this machine's own key, which ssh-keygen -lf /etc/ssh/ssh_host_ed25519_key.pub prints for comparison. (The lab recorded it beforehand, so the terminals below show no prompt.)
Host backup01HostName 127.0.0.1User ca-backupIdentityFile ~/.ssh/backup01_ed25519
On a real network, HostName would be the backup server's name or address. In the command, -a (archive) copies directories recursively and keeps permissions, modification times and symbolic links; owners are kept only when the receiving side runs as root. -v lists each file it sends, and backup01:webapp/ is a path relative to the remote account's home directory.
The first copy sends everything. Run it again, then once more after adding a line to the configuration file:
The second run sends nothing: rsync compares each file's size and modification time and skips the ones that match. The third sends only webapp.conf, and for a large file with a small change it would send only the changed parts. The speedup figure is the total size divided by the bytes actually sent and received. That is why rsync is the tool for repeated copies and deployments. A copy is not yet a backup, though: the next run faithfully copies an accidental deletion or a corrupted file over the good one, so keep older versions too (dated archives, or a backup tool), and back up a running database with its own dump or snapshot tool, not by copying its files while it writes them.
Two details cause most rsync accidents. A trailing slash on the source means "the contents of this directory"; without it, rsync copies the directory itself into the destination. -n (--dry-run) shows what would happen without changing anything:
Without the slash, the files would have landed in webapp/webapp/ on the backup. The second detail is --delete, which makes the destination an exact mirror by deleting files that no longer exist on the source. Aimed at the wrong directory, it deletes the wrong files, so run it with -n first:
The dry run announces deleting logs/error.log, and the backup still has the file because nothing was changed. For a single file, scp FILE backup01: also works. Since OpenSSH 9.0 it transfers over the SFTP protocol; -O selects the legacy SCP protocol for old servers that lack SFTP.
Downloading a file and proving it is genuine
curl downloads a URL. For files, combine -f (fail on an HTTP error instead of saving the server's error page under the file's name), -sS (no progress meter, but still print errors), -L (follow redirects) and -O (save under the name at the end of the URL). wget URL does much the same. Watch one trap between them: curl's -O takes no argument, while wget -O NAME saves under the name you give it.
The number in parentheses, 22, is also curl's exit status, and ls shows that no file was saved. A script that checks the exit status of curl stops here instead of unpacking an HTML error page.
A checksum proves that a file is intact. SHA-256 turns any file into a 64-character fingerprint, and changing a single byte changes the fingerprint completely. Publishers list the fingerprints of their files in a checksum file. The example downloads a small file from the Ubuntu cloud image release this course's lab machines were built from, the image manifest (the list of packages in the image), along with the release's SHA256SUMS and its signature. The variable base only saves typing the long address three times.
The two fingerprints are identical; the * in the published list marks binary mode and changes nothing for the check. sha256sum -c does the comparison for every file named in SHA256SUMS, and --ignore-missing skips the many files you did not download. Change one line of the manifest and the check fails, with exit status 1:
A checksum only tells you that the file matches the list. Someone who controls the download server can replace the file and the list together. A signature closes that gap. Canonical signs SHA256SUMS with its cloud image key, and the public half of that key came with your server in /usr/share/keyrings, from the ubuntu-keyring package. gpgv checks a signature against keys you already have:
Good signature means SHA256SUMS is exactly what the key's owner signed. The forged list, made by adding the changed manifest's own checksum, fails with BAD signature and exit status 1. So the order is: check the signature on the list, then check the file against the list. Package repositories apply the same chain to every package they serve, which is one reason to prefer them to loose downloads.
curl -fsSL URL | sudo bash run whatever the server sends, as root, before anyone has read or verified it, and if the connection drops, bash runs whatever part arrived. Download the script to a file, verify it, read it, and only then run it. Better still, install from the vendor's signed package repository, as the next lesson shows.Try this
Using the backup01 account from this lesson, make a checksummed backup and prove it survived the trip. Create webapp.tar.gz, write its checksum with sha256sum webapp.tar.gz > webapp.tar.gz.sha256, copy both with rsync -a webapp.tar.gz webapp.tar.gz.sha256 backup01:, then run ssh backup01 sha256sum -c webapp.tar.gz.sha256, which should print webapp.tar.gz: OK. Extract it on the other side into a new, empty directory and list it, all in one remote command: ssh backup01 "mkdir -p check && tar -xf webapp.tar.gz -C check && ls check/webapp". Finally, delete a local file and run rsync -avn --delete to read the deleting line before you decide whether the real run is what you want. To clean up, delete the backup01 block from ~/.ssh/config and run:
Takeaway
List an archive before you extract it, and extract it as yourself into an empty directory. Verify a download against a signed checksum before you unpack or run it, and dry-run every rsync that deletes.