Linux interview questions
Prep Linux interviews from file modes and signals through systemd, cgroups v2, and production troubleshooting — tagged Beginner to Expert.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Files & permissions
Each file has read/write/execute bits for owner, group, and others — like 640 for rw-r-----. On directories, execute means you can enter or traverse; read means you can list names.
ls -l app.sh # -rwxr-x--- → user rwx, group r-x, other --- = 750 chmod 640 app.sh chmod u+x,o-r app.shLink to this question
Numeric encodes rwx as bits: r=4, w=2, x=1, summed per class — 755 is rwxr-xr-x. Symbolic edits relative bits like u+x or go-w, which is safer when I only want one change.
chmod 644 file chmod 755 dir chmod g+w,o-rwx fileLink to this question
A hard link is another directory entry for the same inode — same data, same filesystem. A symlink is a path pointer that can cross filesystems and break if the target moves.
ln file hard ln -s file soft ls -li file hard softLink to this question
It's a mask whose bits are cleared from the requested mode (mode & ~umask) — 022 yields 644 files and 755 directories. It sets the default posture for new files per shell or service.
umask touch new; ls -l new # 666 masked by 022 → 644Link to this question
SUID runs a binary as the file owner — like passwd. SGID runs as the group, or makes new files inherit a directory's group. Sticky on a directory like /tmp means only the owner can delete their own files.
SUID and SGID binaries are a classic local privilege-escalation path and belong on a regular audit with find -perm -4000/-2000. Package-managed SUID like passwd or sudo is expected; random SUID under /home or /tmp is not. SGID directories help shared team drops inherit group ownership; sticky on /tmp prevents users from deleting each other's files. A capital S or T in ls means the special bit is set without execute — usually a mistake.
ls -l /usr/bin/passwd # -rwsr-xr-x → s in owner exec = SUID find / -perm -4000 -type f 2>/dev/null | head
Interviewer often follows with: How do file capabilities change the need for SUID?
Link to this questionWhen owner/group/other isn't enough, ACLs add per-user or per-group entries via setfacl and getfacl. A trailing + in ls -l marks that an ACL is present.
setfacl -m u:alice:rw report.txt getfacl report.txt ls -l report.txt # -rw-rw-r--+ → the + means ACLLink to this question
I'd treat it as a critical privilege-escalation risk. Remove SUID or fix ownership and mode, find who created it, and check whether it was already abused.
SUID root plus writable by others is still critical: Linux clears the SUID bit when a user without CAP_FSETID modifies the file, but the tampered binary stays root-owned, so anything root runs it from, such as cron, a unit, or an admin, executes attacker code. Immediate actions for me: chmod u-s, quarantine the path, compare against package manager file integrity, inspect shell history and auth logs, and hunt for unexpected root-owned cron or systemd units. Capital S in ls means the SUID bit is set but execute isn't — usually a misconfiguration. Regular find audits for -4000/-2000 should be part of baseline hardening.
ls -l /usr/local/bin/suspect chmod u-s /usr/local/bin/suspect rpm -Vf /usr/local/bin/suspect 2>/dev/null || debsums -c 2>/dev/null | head find / -perm -4000 -type f 2>/dev/null
Interviewer often follows with: How do capabilities replace many classic SUID needs?
Link to this questionCapabilities split root into discrete privileges — CAP_NET_BIND_SERVICE, CAP_SYS_ADMIN, and so on. I'd grant only what a service needs instead of UID 0 with the full set.
Historically everything needed root; capabilities — and ambient, inherited, bounding sets — let you bind low ports or use raw sockets without a full root shell. systemd can set AmbientCapabilities= and CapabilityBoundingSet=. Containers drop capabilities the same way. CAP_SYS_ADMIN is dangerously broad — I'd treat it nearly like root. File capabilities via setcap can replace SUID for specific binaries, but they're still privilege and must be inventoried. getpcaps and capsh --print help inspect the current set.
capsh --print getpcaps $$ sudo setcap 'cap_net_bind_service=+ep' /usr/local/bin/api getcap /usr/local/bin/api
Interviewer often follows with: What is the difference between permitted, effective, and bounding sets?
Link to this questionI'd own by a deploy group, mode 750/640, optional ACLs for break-glass readers, no world access, and sticky or immutable bits only where justified. Prefer group ownership over chmod 777.
Pattern I'd use: root or deploy owns the tree, group app can read/execute, others get nothing. Shared drop directories use SGID so new files inherit the group, plus sticky if users must not delete each other's files. Secrets stay 600 root:root or a dedicated secrets group, never in world-readable config. ACLs help when two teams need different access without exploding group membership — but document them because ls alone under-communicates. Pair with systemd UMask=, ProtectSystem=, and ReadWritePaths= so the service can't wander the filesystem even if compromised.
chown -R root:app /opt/app
find /opt/app -type d -exec chmod 750 {} \;
find /opt/app -type f -exec chmod 640 {} \;
chmod 750 /opt/app/bin/*
setfacl -m g:audit:rx /opt/app/logsInterviewer often follows with: How would systemd ProtectHome and ProtectSystem complement this?
Link to this questionOnly root can chown arbitrarily. Ordinary users can chmod their own files and may chgrp to a group they belong to, depending on policy.
sudo chown app:app /var/lib/app/data chgrp app report.txt ls -l report.txtLink to this question
Processes & signals
A process has its own address space; threads share one process's memory. Both are kernel tasks — threads just share more via clone() flags.
ps -eLf | head cat /proc/self/status | grep ThreadsLink to this question
SIGTERM asks a process to exit and can be caught for cleanup. SIGKILL can't be caught or ignored. SIGHUP often means "reload config" for daemons.
kill -TERM 4823 kill -HUP 4823 kill -9 4823 kill -lLink to this question
A zombie has exited but its parent hasn't wait()ed — it only holds a PID slot. An orphan's parent died; init or systemd adopts and reaps it.
ps -eo pid,ppid,stat,cmd | awk '$3 ~ /Z/' # fix the parent — you cannot kill a zombie itselfLink to this question
nice from -20 to 19 hints CPU scheduling priority — higher nice is nicer, less CPU. ionice sets I/O priority classes. I'd use them for batch jobs; hard isolation still needs cgroups.
nice is a scheduler hint, not a hard cap — a niced process can still saturate CPUs if nothing else competes. ionice classes only take effect under an I/O scheduler that supports priorities (bfq or mq-deadline); under none or kyber they do nothing. For isolation guarantees I'd use cgroup cpu.max, io.weight, or io.max — systemd CPUQuota=, IOWeight=, IOReadBandwidthMax=. Still use nice/ionice for best-effort batch jobs on shared bastions.
nice -n 10 ./batch.sh ionice -c2 -n7 ./batch.sh renice 15 -p <pid>
Interviewer often follows with: When would you choose cgroup CPUQuota over nice?
Link to this questionA unit file declares ExecStart, dependencies, restart policy, and resource limits. systemd tracks the service as a cgroup, captures logs to the journal, and orders boot via targets.
Each service gets a cgroup so forked children can't escape supervision. Restart= and StartLimitBurst prevent crash loops from hammering the box. I'd use Type=notify or notify-reload when the app supports sd_notify, otherwise Type=exec; systemd discourages Type=forking with a PIDFile. Logs go to the journal. Hardening directives — ProtectSystem, PrivateTmp, CapabilityBoundingSet, NoNewPrivileges — close a lot of post-compromise paths without rewriting the app.
# /etc/systemd/system/app.service [Service] ExecStart=/usr/bin/app Restart=on-failure MemoryMax=512M systemctl enable --now app journalctl -u app -f
Interviewer often follows with: What is the difference between Requires= and Wants=?
Link to this questionI'd confirm what the unit's main PID actually is, whether it traps TERM, and what KillMode= and TimeoutStopSec= do. Escalate to SIGKILL only after I understand cleanup needs.
With the default KillMode=control-group, systemd sends SIGTERM to every process in the unit's cgroup, waits TimeoutStopSec, then sends SIGKILL; mixed and process send TERM to the main PID only. If ExecStart is a shell script without exec, the shell becomes the main PID, so under mixed or process the real server never sees TERM. I've hit that more than once. Fix with exec in the script, a correct Type=, and the default control-group kill mode. strace -p on the main PID during stop shows whether TERM arrives. Containers have the same class of bug when ENTRYPOINT is shell-form.
systemctl cat app systemctl show app -p MainPID -p TimeoutStopUSec -p KillMode strace -p "$(systemctl show -p MainPID --value app)" -e signal journalctl -u app -b --no-pager | tail
Interviewer often follows with: What does KillMode=control-group change vs process?
Link to this questionrescue.target and emergency.target are for broken boots or maintenance — fewer services, root shell. isolate replaces the current target; I'd use it carefully because it stops everything not in the destination.
multi-user.target is the usual server default; graphical.target adds a display manager. rescue is single-user-ish with local FS mounted; emergency is even more minimal. From GRUB you can append systemd.unit=rescue.target. isolate multi-user.target is how you leave rescue. I'd avoid casual isolate on production hosts over SSH — you can kill the session's dependencies.
systemctl get-default systemctl list-units --type=target # GRUB: systemd.unit=rescue.target
Interviewer often follows with: How do you find which unit failed and blocked boot?
Link to this questionCatch SIGTERM/SIGINT for graceful drain, SIGHUP for reload if supported, handle SIGPIPE or EPIPE, and leave SIGKILL to the supervisor as last resort. Document the drain timeout for orchestrators.
Graceful stop means stop accepting work, finish in-flight requests within a deadline, flush logs, then exit 0. systemd TimeoutStopSec and Kubernetes terminationGracePeriodSeconds must be at least that deadline. Reload via SIGHUP shouldn't drop connections when possible. Double-signal patterns — TERM then wait then KILL — are normal. I wouldn't rely on SIGKILL for routine deploys — it skips cleanup and can corrupt on-disk state. Pair with readiness probes so traffic leaves before TERM.
[Service] ExecStart=/usr/bin/app ExecReload=/bin/kill -HUP $MAINPID TimeoutStopSec=30 KillSignal=SIGTERM FinalKillSignal=SIGKILL
Interviewer often follows with: How do you test graceful drain in CI without a full cluster?
Link to this questionI'd identify it with top or ps sorted by CPU or memory, confirm it's safe to stop, send SIGTERM first, then SIGKILL only if it ignores me.
ps -eo pid,%cpu,%mem,cmd --sort=-%cpu | head kill -TERM <pid> kill -9 <pid>Link to this question
PID 1 — systemd on most servers — parents orphans, reaps zombies, and only receives signals it has installed handlers for. Inside containers your app is often PID 1 — it must handle SIGTERM and reap children, or use --init.
docker run --init app:1 # or install tini as ENTRYPOINTLink to this question
Performance & troubleshooting
I'd check load and the four resources: CPU, memory, disk I/O, and network. Correlate with recent changes and journalctl before deep-diving one process.
uptime top -o %CPU free -m df -h journalctl -p err -b --no-pager | tailLink to this question
Average runnable plus uninterruptible — D-state, often I/O — tasks over 1/5/15 minutes. Compare to CPU count — load around cores means busy; sustained much higher means saturation.
uptime nproc # load 3.9 on 4 cores ≈ fully busy, not necessarily overloadedLink to this question
strace shows syscalls — where a process blocks or fails. lsof lists open files and sockets. Together they answer "what is it waiting on?" and "what is holding this file?"
strace answers syscall-level "what is it doing or failing on" — I'd attach with care on latency-sensitive processes and use -e to filter. lsof maps fds to paths and sockets; +L1 finds deleted-but-open files filling disks. For CPU hotspots I'd prefer perf over blind strace. In containers, nsenter or kubectl debug may be required to see the right namespaces. Always confirm you have the right PID — container PID vs host PID.
strace -f -e trace=openat,connect,network -p <pid> lsof -p <pid> lsof +L1 | grep deleted
Interviewer often follows with: How do you strace a process inside a container from the host?
Link to this questionjournalctl -u unit -b shows this boot's logs; -f follows; --since and -p filter time and priority. Persistence follows Storage= in journald.conf: auto keeps logs only if /var/log/journal exists, and upstream systemd 259+ defaults to persistent.
journalctl -u nginx -b --no-pager journalctl -u nginx --since "1 hour ago" -p warning journalctl -k -b | grep -i oomLink to this question
The unified hierarchy that limits and accounts CPU, memory, I/O, and PIDs per leaf. systemd and containers place each service or pod in a cgroup and enforce memory.max / cpu.max.
cgroup v2 is a single unified hierarchy — no split cpu vs memory trees. Controllers like memory, cpu, io, pids enforce limits and accounting per leaf. systemd places units under system.slice; containers get their own leaves so one workload's memory.max OOM doesn't necessarily take the host. Delegation matters for rootless and nested runtimes. systemd 258 removed cgroup v1 support; stat -fc %T /sys/fs/cgroup prints cgroup2fs on a v2 host. When limits disagree — app heap bigger than cgroup memory.max — expect exit 137.
systemctl status app ls /sys/fs/cgroup/system.slice/app.service/ cat /sys/fs/cgroup/system.slice/app.service/memory.current
Interviewer often follows with: Where do you read memory.current for a systemd unit?
Link to this questionA deleted file still open by a process keeps blocks until the fd closes. df counts them; du walking the tree can't see them. I'd find it with lsof and restart or truncate the holder.
Classic with long-lived apps that rotate logs incorrectly — delete instead of truncate or copytruncate. Containers and journald can show the same symptom. Fix: identify deleted-but-open paths via lsof +L1 or /proc/<pid>/fd, truncate the fd or restart the process, then fix log rotation. Also check sparse files and mount bind overlays when numbers still look wrong.
df -h /var lsof +L1 | grep deleted # truncate without restart if safe: # : > /proc/<pid>/fd/3
Interviewer often follows with: How should logrotate be configured to avoid this?
Link to this questionI'd identify the process with top or pidstat, split user vs system time, sample stacks with perf, and check for runaway threads, livelocks, or noisy neighbors in the same cgroup or CPU set.
USE method: utilization, saturation, errors per resource. High %us points at application code; high %sy at syscalls, spinlocks, or networking; high %steal on hypervisors. Softirq overload shows in /proc/softirqs and can look like "CPU wait" without a busy user process. For Java/.NET I'd use async profilers; for native, perf top or flamegraphs. Fix may be code, reducing parallelism, CPU affinity or cgroup cpu.max, or moving off a noisy host. Always capture a short timeline: when it started, deploys, and traffic deltas.
top -o %CPU pidstat -u 1 5 perf top -p <pid> mpstat -P ALL 1 5
Interviewer often follows with: What does sustained high load with low %CPU usually imply?
Link to this questionLinux uses free RAM for page cache and reclaim under pressure. "Low free" is normal; I'd watch available, PSI, swap-in, and thrashing — not just the free column.
free -m available estimates what can be given to new workloads without swapping. Under cgroup limits, a container can OOM while the host still has cache. PSI files under /proc/pressure/ show some/full stalls for cpu/memory/io and are excellent early warnings. Dropping caches with drop_caches is a diagnostic hammer, not a fix — it can hurt performance. Prefer fixing the consumer, sizing limits, and making sure working sets fit.
free -m cat /proc/pressure/memory vmstat 1 5 sar -r 1 5
Interviewer often follows with: Why can a container OOM while free -m on the host still looks fine?
Link to this questionWhen I need cheap, system-wide answers that sampling tools miss: short-lived processes, per-I/O disk latency, TCP retransmits, or who keeps opening a file. execsnoop, biolatency, and tcpretrans answer those in seconds without the ptrace overhead strace adds.
strace works through ptrace and can slow a busy process badly; eBPF programs run in the kernel and aggregate there, so summaries like biolatency or runqlat stay cheap under load. Typical wins: execsnoop for crash-looping helpers that top never catches, opensnoop for config and permission hunts, biosnoop and biolatency for disk latency by process, tcpconnect and tcpretrans for network flakiness, offcputime for where threads block. bpftrace one-liners cover the rest. They need root, or CAP_BPF plus CAP_PERFMON on Linux 5.8+, so treat access to them like root on production hosts.
execsnoop # bcc: every exec, even sub-second ones
biolatency 1 10 # bcc: block I/O latency histogram, 1 s x 10
bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @[comm] = count(); }'Interviewer often follows with: Why is eBPF tracing access on a production host close to root access?
Link to this questionNetworking & boot
ss -tulpn — or lsof -i :PORT — shows the listener, PID, and process. I'd trace that PID back to a systemd unit or container.
ss -tulpn | grep :8080 ps -p <pid> -o pid,cmd systemctl status <pid>Link to this question
Client SYN → server SYN-ACK → client ACK → ESTABLISHED. Close is a FIN/ACK exchange; the active closer usually enters TIME_WAIT to absorb late packets.
ss -tan state established | head ss -tan state time-wait | wc -lLink to this question
nsswitch.conf orders sources — files, DNS, and so on. Apps call getaddrinfo; systemd-resolved often sits in the middle. dig talks to DNS directly — getent shows what apps see.
getent hosts api.internal resolvectl query api.internal cat /etc/nsswitch.conf | grep hostsLink to this question
ss answers "what sockets exist and in what state". tcpdump answers "what packets are on the wire". I'd use ss first, then tcpdump when I need handshake or payload-level proof.
ss -tulpn tcpdump -ni eth0 port 443 -c 20 tcpdump -ni any host 10.0.0.5 and port 5432 -c 50Link to this question
I'd resolve the name, ping or arp the IP, check the route, then test the port and firewall. Localize whether it's DNS, L3, or L4/policy.
A structured path beats random tool spam: getent or resolvectl for DNS, ip route get for egress path and interface, ping only if ICMP is allowed — failure is inconclusive on filtered networks — traceroute or mtr for path, then nc or curl for the port. Check nftables/iptables and security groups in parallel. On servers, also verify the service is bound to the expected address — 0.0.0.0 vs 127.0.0.1.
getent hosts app.example ip route get 10.0.0.5 ping -c1 10.0.0.5 nc -vz 10.0.0.5 443 nft list ruleset | head
Interviewer often follows with: Why can ping fail while HTTPS works?
Link to this questionPackets hit hooks — prerouting, input, forward, output, postrouting. Chains match and accept, drop, reject, or jump. nftables is the modern unified framework; iptables often translates into it.
Local delivery uses input; routed traffic uses forward. Docker and kube-proxy historically inserted many iptables rules — order and firewalld coexistence cause surprising drops. I'd always list counters to see which rule matches. Prefer one firewall manager; mixing ufw, firewalld, and raw iptables invites last-change-wins bugs. Conntrack state matters for established flows and NAT.
nft list ruleset iptables -L -n -v --line-numbers iptables -t nat -L -n -v
Interviewer often follows with: Where would you look for Docker-published port DNAT rules?
Link to this questionUEFI → GRUB → kernel plus initramfs → systemd as PID 1 → targets and units → getty. Each stage has distinct logs; the last message tells you where it stuck.
Firmware/POST failures never reach GRUB. GRUB issues point at bootloader config or disk. initramfs hangs often mean missing modules, wrong root= UUID, or encrypt/unlock prompts. After switch-root, systemd records unit failures; journalctl -b -p err and systemctl --failed are my first stops. For early boot, enable persistent journal or read /run/log when disk wasn't writable. Kernel panics dump to console; remote serial or IPMI helps headless hosts. Masking a bad unit from a live USB is a common recovery path.
systemctl --failed journalctl -b -p err --no-pager | tail -n 50 journalctl -u NetworkManager -b
Interviewer often follows with: How do you boot into rescue from GRUB without a live USB?
Link to this questionI'd capture both ends with tcpdump, compare seq/ack and who sent RST, check conntrack exhaustion, middleboxes, and app idle timeouts. Correlate with load balancer health flaps.
An RST from the local stack often means nothing listening, or an application closed aggressively. RSTs from a middlebox can indicate ACL/IDS or asymmetric routing. TIME_WAIT accumulation and nf_conntrack_table full show up in dmesg and cause drops that look like resets. Idle timeouts on LBs shorter than app keepalives create client-visible blips — fix with better health checks and idle settings. Always note source ports and whether only one AZ or path is affected.
tcpdump -ni eth0 host 10.0.0.8 and port 5432 -w /tmp/db.pcap ss -s dmesg | grep -i conntrack conntrack -C 2>/dev/null
Interviewer often follows with: How do you tell an application RST from a firewall RST in a pcap?
Link to this questionINPUT filters traffic destined for the local host. FORWARD filters traffic routed through the host — containers, VMs, routers. Misplacing rules is a common Docker/firewall footgun.
iptables -L INPUT -n -v iptables -L FORWARD -n -v # bridge/container traffic often needs FORWARD allowLink to this question
ip route show default for the gateway and iface; resolvectl status or /etc/resolv.conf for resolvers. Wrong route or empty resolvers look like "the network is down."
ip route show default ip -br addr resolvectl status | headLink to this question
Real-world scenarios
OOM score is heuristic — unprotected critical processes can die first. I'd raise protection for sshd and agents, put the batch job in a memory cgroup with a hard limit, and fix the leak.
The killer picks the process with the highest oom_score, which is driven by memory footprint and shifted by oom_score_adj; process age has not counted since 2.6.36. Large, unprotected allocators lose; processes with oom_score_adj=-1000 are immune — careful: mistaking a leaky app for "critical" just panics the box later. systemd has OOMScoreAdjust=; containers need memory.max so the workload dies inside its cgroup before the host. Post-incident: journalctl -k or dmesg for the kill, inspect /proc/<pid>/oom_score, and confirm whether it was a global OOM or a cgroup OOM. In cgroup v2, hitting memory.max often kills inside that cgroup first (a container dies with 137), and memory.events oom_kill counts kills per cgroup; a host-wide OOM means the machine was overcommitted. On hosts running systemd-oomd, a userspace kill of a whole cgroup is logged by systemd-oomd, not as a kernel OOM line. Long term: limits on batch queues, JVM/Node heaps sized to fit their cgroup, swap policy awareness, and alerts on memory PSI (/proc/pressure/memory or the cgroup's memory.pressure).
journalctl -k -b | grep -i 'killed process' cat /proc/$(pgrep -o sshd)/oom_score /proc/$(pgrep -o sshd)/oom_score_adj cat /sys/fs/cgroup/batch.slice/memory.events # oom_kill count for the slice # systemd drop-in: OOMScoreAdjust=-500 for sshd; MemoryMax= for the batch slice
Interviewer often follows with: How does a cgroup OOM differ from a system-wide OOM in symptoms?
Link to this questionI'd stagger service starts with randomized delay and dependencies, shed non-critical unit Wants, and add client-side backoff with jitter so reconnects don't align.
Cold boot aligns cron, systemd After=network-online, and app reconnect loops. Fix layers: systemd RestartSteps=/RestartMaxDelaySec= for growing backoff (254+) plus RestartRandomizedDelaySec= jitter (262+), rate-limited timers, load-balancer slow-start, and server-side admission control. For NFS, enable graceful reconnect and avoid simultaneous fsck/mount storms across the fleet. Runbooks should include "boot into rescue and start tiers manually" when automation makes the outage worse. I'd chaos-test reboot of N% of the fleet, not only single hosts.
systemctl show myapp | grep -E 'Restart|After' # drop-in: RestartSec=30, RandomizedDelaySec= on timers # app: exponential backoff + jitter on DB connect
Interviewer often follows with: Where do you put jitter — systemd, the app client, or the load balancer?
Link to this questionInodes. df -i often shows 100% while block usage looks fine — millions of tiny files exhausted the inode table.
Inode exhaustion is classic on ext4, where the inode count is fixed at mkfs, with mail spools, container overlay trees, or app session caches; XFS allocates inodes dynamically up to imaxpct, which xfs_growfs -m can raise. Diagnosis: df -i, find directories with huge file counts, identify the producer. Remediation: delete or archive tiny-file trees, raise inode ratio only on mkfs — too late for a live FS — and move workloads to filesystems sized for the file-count pattern. Also distinguish quota failures from true ENOSPC. Monitor both block and inode usage.
df -h / df -i / find /var/spool -xdev -type f | wc -l du -sh /var/spool/* 2>/dev/null | sort -h | tail
Interviewer often follows with: Can you add inodes to an existing ext4 filesystem without recreating it?
Link to this questionI'd verify the listener, cert chain and clock, then capture the handshake. Usual causes: expired cert, SNI mismatch, missing intermediate, or TLS version/cipher mismatch.
openssl s_client -connect host:443 -servername … shows chain, dates, and alert codes. Check system time — skew breaks validity — file permissions on key material, and whether the process reloaded after cert rotation. tcpdump helps spot RST mid-handshake vs alert. Intermediate-not-served is still common with incomplete fullchain.pem. For mTLS, verify client cert CA trust on the server side separately. Application "connection reset" logs often hide handshake alerts — I'd always get the TLS layer view.
date -u openssl s_client -connect api.internal:443 -servername api.internal </dev/null 2>&1 | openssl x509 -noout -dates -subject ss -tulpn | grep ':443'
Interviewer often follows with: How do you tell a missing intermediate from a hostname mismatch in s_client output?
Link to this questionTasks are blocked on storage in D state, and Linux load average counts those alongside runnable tasks, so load looks huge while CPUs sit idle waiting. I'd find the hot device and the processes stuck in D state (iowait alone is an unreliable signal, so confirm with per-device await and %util), then decide whether it's cold-cache reads, a write storm, or a failing disk.
Load average counts runnable plus uninterruptible sleep. Heavy NFS, failing disks, or sync-heavy writes produce D-state piles and high wa% without busy CPU. iowait means CPUs are idle waiting on I/O, and storage latency is saturation even when CPU looks idle. Tools: iostat -xz for await and %util on the busy device, pidstat -d for the writers, ps for wchan, nfsstat for network mounts, and dmesg for storage errors. Fix the I/O path (faster volume, less fsync amplification, repaired NFS, isolated noisy neighbors); adding CPU cores won't help, and parallelism that stampedes the same device makes it worse. Distinguish steal% from iowait. Container hosts often hide the writer as a container PID; map back with nsenter or docker top.
uptime; mpstat 1 5 iostat -xz 1 5 pidstat -d 1 5 ps -eo pid,stat,wchan:32,cmd | awk '$2 ~ /D/'
Interviewer often follows with: Why can NFS latency inflate load average more than local SSD latency for the same app?
Link to this questionI'd stop the restart storm — mask or set StartLimit — read the first failure in the journal, fix the Exec or config, then re-enable with sane RestartSec limits.
Restart=always with a RestartSec= long enough to stay under the default start limit (5 starts per 10 s) turns a config typo into a noisy loop that fills the journal and can thundering-herd dependencies. Triage: systemctl status, journalctl -u -b, look at ExitCode/StatusErrno, run the ExecStart by hand under the same User= and Environment=. Temporary: systemctl mask --now or edit Restart=on-failure with StartLimitIntervalSec=. Permanent: fix the binary path, permissions, or After= ordering. Avoid RemainAfterExit mistakes for Type=oneshot services that look "active" while broken.
systemctl status myapp --no-pager journalctl -u myapp -b -p err --no-pager | head -n 40 systemctl mask --now myapp # stop the storm while fixing # drop-in: [Unit] StartLimitIntervalSec=300 StartLimitBurst=3; [Service] RestartSec=10
Interviewer often follows with: What is the difference between systemctl disable and systemctl mask during an incident?
The bounding set is a hard ceiling: nothing the unit starts, including a binary with file capabilities or a helper that re-execs, can gain a capability outside it. So if CAP_SYS_ADMIN survives, the effective CapabilityBoundingSet= still contains it. I'd read the merged value with systemctl show and then check the live process.
CapabilityBoundingSet= lines are merged: plain lines are ORed together, a line prefixed with ~ removes the listed capabilities, an empty assignment (often in a drop-in) resets the set, and a bare ~ restores the full set. Commands prefixed with + in ExecStart= are not bounded by it. The service must be restarted after daemon-reload before a running process changes, and a process started outside the unit (a cron job, docker exec, a sidecar) is not bounded by it at all. AmbientCapabilities= cannot exceed the bounding set, so it is not the gap. Verify with systemctl show -p CapabilityBoundingSet, getpcaps on the live PID and grep Cap /proc/PID/status, and treat CAP_SYS_ADMIN as near-root.
systemctl show -p CapabilityBoundingSet myapp # merged result of every line systemctl cat myapp # look for ~ lines and empty resets in drop-ins pid=$(systemctl show -p MainPID --value myapp); getpcaps $pid grep Cap /proc/$pid/status
Interviewer often follows with: Does NoNewPrivileges= block ambient capabilities already granted to the service?
Link to this questionI'd compare /proc/<pid>/ns/* inodes to the host init namespace. Matching net or pid inodes means the container was started with host namespaces or joined them.
Each namespace has an inode; ls -l /proc/1/ns/net vs /proc/<container_pid>/ns/net tells you if they share. docker inspect HostConfig.NetworkMode/PidMode/UTSMode should match. Causes: --network=host, --pid=host, mis-set RuntimeClass, or a breakout that joined namespaces. Impact: host network sockets and process signalling become reachable — treat as isolation failure. Remediate by redeploying with private namespaces, then hunt for how the flag was introduced. Pair with runtime detections on setns and host-ns starts.
ls -l /proc/1/ns/net /proc/1/ns/pid
pid=$(docker inspect -f '{{.State.Pid}}' app)
ls -l /proc/$pid/ns/net /proc/$pid/ns/pid
docker inspect app --format 'Net={{.HostConfig.NetworkMode}} Pid={{.HostConfig.PidMode}}'Interviewer often follows with: Can two containers share a netns with each other without sharing the host netns?
Link to this questionI'd fix time sources first — chrony or ntp — verify step vs slew, then restart time-sensitive services. I wouldn't mass-rotate certs until clocks are sane.
Skew breaks Kerberos tickets, TLS notBefore/notAfter, and cookie expirations. Recovery: check timedatectl/chronyc tracking, restore reliable NTP peers, allow chrony to step if far skewed, then bounce SSSD/httpd/app pools. Avoid generating new certs "because TLS failed" while clocks are wrong — you'll mint more confusion. Prevent with multiple NTP sources, monitoring offset, and guest VM tools sync on hypervisors. Document that auth outages can be time, not IdP.
timedatectl status chronyc tracking chronyc makestep # then systemctl restart sssd httpd
Interviewer often follows with: When is stepping the clock safer than slewing during an auth outage?
Link to this questionI'd count FDs under /proc/<pid>/fd, raise the limit only as a bridge, then find the leak — sockets, files, or never-closed JDBC — and patch or restart with a hard ulimit ceiling.
Diagnosis: ls /proc/pid/fd | wc -l vs LimitNOFILE from systemctl show, ss -s for socket accumulation, lsof -p for the dominant type. Soft limits can be raised live carefully; hard limits need unit edits. Temporary: restart to reclaim, scale out, or kill runaway children. Root cause: connection pool mis-size, FD leak on exception paths, or log file handles. Set LimitNOFILE= deliberately — unlimited hides leaks until the host fails. Alert on fd usage percent of the limit.
pid=$(pgrep -o java); ls /proc/$pid/fd | wc -l systemctl show myapp | grep LimitNOFILE ss -s # drop-in: LimitNOFILE=65535 — then fix the leak
Interviewer often follows with: How do you tell a connection-pool misconfiguration from a true FD leak?
Link to this questionI'd check for AVC denials in the audit log, restore contexts, and only then consider a focused boolean — I wouldn't chmod 777 to "make it work."
RPM updates can leave mislabeled files if admins copied trees without restorecon. Diagnosis: ausearch -m AVC -ts recent, sealert, ls -Z vs expected type. Fix: restorecon -Rv on the tree, ensure the unit uses the right SELinux domain, or ship a proper policy module for custom paths. setenforce 0 is a diagnostic toggle, not a production fix. On AppArmor hosts the analog is dmesg/journal DENIED plus aa-status. Interview signal: label/MAC before blaming application code.
ausearch -m AVC -ts recent | tail ls -Z /var/www/app restorecon -Rv /var/www/app getenforce
Interviewer often follows with: How do you permanently allow a custom data directory without disabling SELinux?
Link to this questionI'd treat it as reclaim/scheduler distress: check memory PSI, thrashing, I/O in D state, and recent kernel/driver changes — not only user CPU profiles.
Soft lockup messages mean a CPU failed to schedule within the threshold — often while stuck in kernel reclaim, a bad driver, or holding a lock under memory pressure. Collect: dmesg/journal -k around the event, /proc/pressure/*, vmstat si/so, and whether transparent huge pages or a specific filesystem ioctl correlates. Mitigations: reduce memory overcommit, tune oom/cgroup limits so a single tenant dies first, update kernel for known reclaim bugs, and capture a sysrq dump if hangs reproduce. I wouldn't "just add swap" as the only fix — it can lengthen thrash windows.
dmesg -T | grep -i 'soft lockup\|hung task' cat /proc/pressure/memory vmstat 1 5 # capture if safe: echo w > /proc/sysrq-trigger (blocked tasks), echo l > /proc/sysrq-trigger (CPU backtraces)
Interviewer often follows with: How does PSI memory pressure change your response compared to classic free -m?
Link to this questionRelated
- Cheat sheetLinux command cheat sheet
- Cheat sheetBash scripting cheat sheet
- Cheat sheetPython for DevSecOps cheat sheet
- CourseLinux essentials
- CourseLinux hardening
- CourseBash for ops, done safely
- Field noteBash strict mode: writing ops scripts that fail loudly
- Field noteSandboxing Linux services with systemd security directives
- Field noteLinux capabilities: dropping root the right way
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.