The CPU scheduler
EEVDF, priorities, real-time and run-queue latency.
When more tasks are runnable than there are CPUs, the scheduler decides which one waits and for how long. This lesson shows how current kernels make that decision, how nice values, cgroup weights and quotas, real-time policies and CPU affinity change it, and how to measure the waiting itself with context-switch counts, schedstat, run-queue latency histograms and CPU pressure. Every experiment crowds two or three tasks onto CPU 1 of a two-vCPU machine, so each effect shows up as a plain percentage. The previous lessons' /proc/PID/schedstat and process states are assumed.
Run queues, classes and EEVDF
Each CPU keeps a run queue of the tasks that are ready to run on it (state R in ps). When the CPU needs its next task, the kernel asks its scheduling classes in a fixed order, and the first class with a runnable task wins. Almost every process runs in the fair class; the classes above it exist for work that must not wait behind it.
Inside the fair class, the kernel tracks each task's virtual runtime: the CPU time it has received, divided by its weight, so a heavier task's clock runs slower. A task's lag is the share it should have had so far minus what it got. Positive (or zero) lag means it is owed CPU and is eligible to run; negative lag means it has had more than its share and waits. Each eligible task also has a virtual deadline: its virtual runtime plus one slice, again scaled by its weight. The scheduler runs the eligible task with the earliest deadline.
An example: two CPU-bound tasks share one CPU, one at the default weight and one about nine times lighter. The light task runs one 1.4 ms slice, but its virtual clock runs nine times faster, so it ends that turn about 13 ms ahead and is not eligible again until the heavy task has run roughly nine slices to catch up. That gives the 90:10 split measured later in this lesson. The rule, Earliest Eligible Virtual Deadline First (EEVDF), has been the fair scheduler since kernel 6.6, so both lab kernels use it; CFS, its predecessor from 2.6.23 until 6.6, simply ran the task with the smallest virtual runtime.
Tuning guides written for older kernels set kernel.sched_latency_ns and kernel.sched_min_granularity_ns. Kernel 5.13 moved them from /proc/sys into debugfs, and 6.6 replaced them with base_slice_ns, so sysctl cannot find them and a file in /etc/sysctl.d that sets them does nothing. The base slice is 0.7 ms multiplied by 1 plus the base-2 logarithm of the CPU count (counting at most 8 CPUs): 1.4 ms on this two-vCPU machine, 2.8 ms on eight or more. RHEL 10 shows the same 1400000. debugfs is readable by root only and is a debugging interface, not a supported setting.
chrt -p shows a task's policy. Policy 0 in /proc/PID/sched is SCHED_OTHER, and prio 120 means nice 0 (the kernel stores 120 plus the nice value). se.load.weight is the nice-0 weight of 1024, multiplied by 1024 for fixed-point arithmetic, and se.slice is the base slice, which util-linux 2.41's chrt also prints as the runtime parameter. Since kernel 6.12 a task can ask for its own slice with the sched_setattr() system call, which chrt --sched-runtime passes; the kernel clamps it to between 0.1 ms and 100 ms. A shorter slice gives earlier virtual deadlines, which the kernel documentation describes as a way for latency-sensitive tasks to be picked sooner; it does not raise their share.
eBPF (BPF) lets small programs, checked by the kernel before they load, run inside the kernel on events and keep counts there; the eBPF tools lesson explains it. Both kernels are built with sched_ext, which lets a scheduler written as BPF programs take over from the fair class. disabled means none is loaded, the default on both distributions. If a loaded BPF scheduler fails or stalls a runnable task, the kernel moves tasks back to the fair class.
nice is a weight
Every fair task has a weight derived from its nice value, from -20 to 19. Nice 0 is 1024, and each step multiplies or divides it by about 1.25, which the kernel describes as roughly 10% of CPU per level; nice 10 is 110. Weights only matter while tasks compete for the same CPU. Two CPU-bound tasks at nice 0 and nice 10 on one CPU should therefore split it 1024 to 110, about 90.3% and 9.7%.
On a two-vCPU machine the tasks must be forced onto one CPU to compete. CPU affinity is the set of CPUs a task may run on: taskset -c 1 CMD limits a command to CPU 1, and stress-ng --taskset 1 does the same for its workers. As in the method lesson, the workload runs as you in a transient scope with guard rails.
pgrep -n picks the newest worker, and renice raised its nice value to 10. The measured split is 90.24% to 9.76%, and each worker's %wait is the other's share: at any moment one runs and one waits. se.load.weight reads 112640 for the reniced worker, 110 times 1024. Nice changes a task's share, not its right to run: if the nice 0 worker stopped, the other would get the whole CPU. You may raise the nice value of your own processes; lowering it again, even back to 0, needs root (CAP_SYS_NICE) or an RLIMIT_NICE allowance.
Between services, cgroup weights decide
When the cgroup cpu controller is enabled, the scheduler works in levels. It first divides CPU time between sibling cgroups according to their cpu.weight (default 100, range 1 to 10000), then between the tasks inside each group according to nice. Here two services each run one worker on CPU 1, and the batch worker is reniced to 19.
The renice changed nothing: 50.20% and 50.00%. Each scope is its own group with weight 100, and nice only ranks tasks inside the batch group, which holds one task. Whether services are separate groups depends on whether systemd has enabled the cpu controller, the part of the cgroup code that divides CPU time, in system.slice. It does that for every unit in the slice as soon as one of them sets a CPU property. A cgroup's cgroup.subtree_control file lists the controllers it passes down to its children:
On a default Ubuntu Server install, the enabled multipathd.service sets CPUWeight=1000, so the controller is on and nice never crosses a service boundary; a host where multipathd is disabled may behave like Rocky below. The Rocky Linux 10 lab machine has no unit like that: system.slice passes no cpu controller down, all services share one group, and the same experiment with nice 10 splits 90 to 10. Check cgroup.subtree_control on your own hosts before relying on nice between services.
Between services, the setting to use is CPUWeight=, which systemd writes to cpu.weight.
Weights of 100 and 20 predict 83.3% and 16.7%; pidstat measured 83.63% and 16.57%. A weight is work-conserving: it only matters while groups compete, so when the web scope stopped, the batch worker at weight 20 took all of CPU 1. A quota is not work-conserving.
CPUQuota=25% became cpu.max 25000 100000: 25 ms of CPU time per 100 ms period, after which the group is throttled until the next period even when the CPU has nothing else to do. CPU 1 is now about 75% idle while the batch job waits. Its time appears as %nice because the worker still runs at nice 19. Use weights to protect one service from another and a quota when a hard ceiling is the requirement; the resource-control lesson covers the systemd settings in depth.
Real-time policies
A SCHED_FIFO task, priority 1 to 99, runs until it blocks, yields or a higher-priority real-time task becomes runnable. Fair tasks get only the time real-time tasks leave, apart from one reservation shown below. SCHED_RR is the same but rotates tasks of equal priority every kernel.sched_rr_timeslice_ms (100 ms). Setting either needs CAP_SYS_NICE, so sudo, unless RLIMIT_RTPRIO allows it. Here a fair worker and a worker switched to SCHED_FIFO share CPU 1.
The FIFO worker took 95.41% and the fair worker kept 4.99%. Older documentation credits that 5% to RT throttling through kernel.sched_rt_runtime_us, 950000 of every 1000000 microseconds. Since kernel 6.12, on kernels built without RT group scheduling (Ubuntu and RHEL both leave it off), that sysctl only limits admission of SCHED_DEADLINE tasks. The 5% now comes from the fair server: a deadline-class reservation on each CPU that runs waiting fair tasks for 50 ms of every second (its runtime and period files in debugfs are in nanoseconds).
With no fair task waiting, the FIFO worker got the whole CPU: the reservation only takes time when fair tasks need it. RHEL 10's 6.12 kernel measured the same pattern: 94.40% and 6.00%, then 100.20% alone. Stop the scope with sudo systemctl stop k-sched-rt.scope.
Measuring run-queue latency
Run-queue latency is the time a task spends runnable but not running: from the moment it wakes up or is preempted until it gets a CPU. The next setup crowds CPU 1 once more with two CPU-bound workers and a probe, stress-ng --cyclic, which sleeps 1 ms in a loop like a service handling small requests (--cyclic-sleep is in nanoseconds). The kernel truncates process names to 15 characters, so the probe shows up as stress-ng-cycli.
pidstat -w counts context switches per second. Each CPU-bound worker is switched out about 610 times a second without asking (nvcswch/s): it is preempted. The probe gives up the CPU about 940 times a second voluntarily (cswch/s), once for every sleep. Mostly nonvoluntary switches mean a task is waiting for a CPU; mostly voluntary ones mean it waits for something else.
The second field of schedstat is the total time in nanoseconds a task waited in the run queue, and the third is how many times it was scheduled. The first worker waited 7.67 s over 9155 turns, about 0.84 ms each. The probe waited 71.5 ms over 14305 wakeups, about 5 microseconds each. That is EEVDF's lag at work: the probe uses a tiny part of its share, so whenever it wakes it is eligible and has the earliest deadline. A histogram shows the whole distribution. On Ubuntu 26.04 the bcc version of the tool fails:
bcc compiles each tool against the kernel headers when it starts, and the bcc 0.35 in the Ubuntu archive fails to compile this tool against the kernel 7.0 headers. bpftrace, installed by default on Ubuntu Server, ships the same tool as runqlat.bt; it reads kernel types from the kernel's built-in BTF data instead of headers. On RHEL 10 the bcc version works as /usr/share/bcc/tools/runqlat. Both run until you press Ctrl-C; timeout -s INT 5 does that after five seconds. The tool runs on every wakeup and context switch on every CPU, and its man page warns that for some workloads its overhead becomes significant: measure it on a test machine first and keep production runs short.
Read the histogram as two populations. The short waits, 1 to 16 microseconds, are the probe and the system's own tasks. A second peak between 0.5 and 2 ms is the two workers waiting for each other's turn, about one base slice of 1.4 ms. An average would have hidden that shape.
perf sched records every scheduler event and perf sched latency summarises them per task. It needs sudo; on Ubuntu, kernel.perf_event_paranoid is 4, which blocks every unprivileged use of perf. The workers waited 0.903 ms on average; the probe waited 0.017 ms on average. The maxima, 16.7 ms and 7.3 ms, are single events in three seconds: an average hides them, and a histogram shows how rare they are. The recording's size grows with the number of CPUs times the switch rate: 2.1 MB for three seconds on this two-vCPU machine, many times that on a large busy host. There, record for one to three seconds with -o on a filesystem with room, and check the output for lost events.
The scope's cpu.pressure gives the same picture as a share of time: in the last ten seconds, at least one of its tasks was waiting for a CPU 94% of the time (some). full, all of its tasks waiting at once, stays near zero because one of the two workers is almost always running. The next step follows from which tasks wait. If they are the ones users wait for, separate them from the competing work with affinity, a cpuset or weights, or reduce that work. If a slow service shows microsecond run-queue delays, the scheduler is not the cause, and the next place to look is locks, I/O or whatever the service calls.
Try this
With the batch and probe scopes still running, move one worker to CPU 0 with taskset -cp 0 $(pgrep -nx stress-ng-cpu) and run pidstat -u -w -C stress-ng-cpu 5 1. Predict first. Expect both workers at about 100%, each on its own CPU, with %wait of a percent or two. The worker still on CPU 1 keeps about 960 nonvoluntary switches a second, because the probe wakes a thousand times a second and preempts it; the moved worker shows fewer than 20. Then stop both scopes with sudo systemctl stop k-sched-batch.scope k-sched-probe.scope and delete the recording with sudo rm /var/tmp/k-sched.perf.
Takeaway
Measure the wait (schedstat, runqlat, CPU pressure) before you change priorities, then choose the lever that matches the boundary: nice inside a service, CPUWeight= between services, a quota for a hard ceiling, affinity to separate work, and real-time policies only for short, bounded work.