The CPU scheduler

EEVDF, priorities, real-time and run-queue latency.

Advanced16 min · lesson 4 of 21

When more tasks are runnable than there are CPUs, the scheduler decides which one waits and for how long. This lesson shows how current kernels make that decision, how nice values, cgroup weights and quotas, real-time policies and CPU affinity change it, and how to measure the waiting itself with context-switch counts, schedstat, run-queue latency histograms and CPU pressure. Every experiment crowds two or three tasks onto CPU 1 of a two-vCPU machine, so each effect shows up as a plain percentage. The previous lessons' /proc/PID/schedstat and process states are assumed.

Run queues, classes and EEVDF

Each CPU keeps a run queue of the tasks that are ready to run on it (state R in ps). When the CPU needs its next task, the kernel asks its scheduling classes in a fixed order, and the first class with a runnable task wins. Almost every process runs in the fair class; the classes above it exist for work that must not wait behind it.

Which class supplies the next task on a CPU
1stop
kernel stopper threads, such as task migration
2deadline (SCHED_DEADLINE)
runtime-per-period reservations
3real-time (SCHED_FIFO, SCHED_RR)
priorities 1 to 99, highest first
4fair (SCHED_OTHER, BATCH, IDLE)
EEVDF: earliest eligible virtual deadline
5ext (SCHED_EXT)
a BPF scheduler, only when one is loaded
6idle
runs when nothing else is runnable
The order is the same on kernel 7.0 (Ubuntu 26.04) and kernel 6.12 (RHEL 10).

Inside the fair class, the kernel tracks each task's virtual runtime: the CPU time it has received, divided by its weight, so a heavier task's clock runs slower. A task's lag is the share it should have had so far minus what it got. Positive (or zero) lag means it is owed CPU and is eligible to run; negative lag means it has had more than its share and waits. Each eligible task also has a virtual deadline: its virtual runtime plus one slice, again scaled by its weight. The scheduler runs the eligible task with the earliest deadline.

An example: two CPU-bound tasks share one CPU, one at the default weight and one about nine times lighter. The light task runs one 1.4 ms slice, but its virtual clock runs nine times faster, so it ends that turn about 13 ms ahead and is not eligible again until the heavy task has run roughly nine slices to catch up. That gives the 90:10 split measured later in this lesson. The rule, Earliest Eligible Virtual Deadline First (EEVDF), has been the fair scheduler since kernel 6.6, so both lab kernels use it; CFS, its predecessor from 2.6.23 until 6.6, simply ran the task with the smallest virtual runtime.

deploy@web01 · Ubuntu 26.04 LTS
$ sysctl kernel.sched_latency_ns kernel.sched_min_granularity_ns
sysctl: cannot stat /proc/sys/kernel/sched_latency_ns: No such file or directory sysctl: cannot stat /proc/sys/kernel/sched_min_granularity_ns: No such file or directory
$ sudo cat /sys/kernel/debug/sched/base_slice_ns
1400000

Tuning guides written for older kernels set kernel.sched_latency_ns and kernel.sched_min_granularity_ns. Kernel 5.13 moved them from /proc/sys into debugfs, and 6.6 replaced them with base_slice_ns, so sysctl cannot find them and a file in /etc/sysctl.d that sets them does nothing. The base slice is 0.7 ms multiplied by 1 plus the base-2 logarithm of the CPU count (counting at most 8 CPUs): 1.4 ms on this two-vCPU machine, 2.8 ms on eight or more. RHEL 10 shows the same 1400000. debugfs is readable by root only and is a debugging interface, not a supported setting.

deploy@web01 · Ubuntu 26.04 LTS
$ chrt -p $$ grep -E '^(policy|prio|se.load.weight|se.slice) ' /proc/$$/sched
pid 251935's current scheduling policy: SCHED_OTHER pid 251935's current scheduling priority: 0 pid 251935's current runtime parameter: 1400000 se.load.weight : 1048576 policy : 0 prio : 120 se.slice : 1400000
$ chrt --other --sched-runtime 100000 0 sh -c 'grep se.slice /proc/$$/sched'
se.slice : 100000

chrt -p shows a task's policy. Policy 0 in /proc/PID/sched is SCHED_OTHER, and prio 120 means nice 0 (the kernel stores 120 plus the nice value). se.load.weight is the nice-0 weight of 1024, multiplied by 1024 for fixed-point arithmetic, and se.slice is the base slice, which util-linux 2.41's chrt also prints as the runtime parameter. Since kernel 6.12 a task can ask for its own slice with the sched_setattr() system call, which chrt --sched-runtime passes; the kernel clamps it to between 0.1 ms and 100 ms. A shorter slice gives earlier virtual deadlines, which the kernel documentation describes as a way for latency-sensitive tasks to be picked sooner; it does not raise their share.

deploy@web01 · Ubuntu 26.04 LTS
$ grep CONFIG_SCHED_CLASS_EXT /boot/config-$(uname -r) cat /sys/kernel/sched_ext/state
CONFIG_SCHED_CLASS_EXT=y disabled

eBPF (BPF) lets small programs, checked by the kernel before they load, run inside the kernel on events and keep counts there; the eBPF tools lesson explains it. Both kernels are built with sched_ext, which lets a scheduler written as BPF programs take over from the fair class. disabled means none is loaded, the default on both distributions. If a loaded BPF scheduler fails or stalls a runnable task, the kernel moves tasks back to the fair class.

nice is a weight

Every fair task has a weight derived from its nice value, from -20 to 19. Nice 0 is 1024, and each step multiplies or divides it by about 1.25, which the kernel describes as roughly 10% of CPU per level; nice 10 is 110. Weights only matter while tasks compete for the same CPU. Two CPU-bound tasks at nice 0 and nice 10 on one CPU should therefore split it 1024 to 110, about 90.3% and 9.7%.

On a two-vCPU machine the tasks must be forced onto one CPU to compete. CPU affinity is the set of CPUs a task may run on: taskset -c 1 CMD limits a command to CPU 1, and stress-ng --taskset 1 does the same for its workers. As in the method lesson, the workload runs as you in a transient scope with guard rails.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=k-sched-nice --uid=$USER -p MemoryMax=256M -p TasksMax=16 \ stress-ng --cpu 2 --taskset 1 --timeout 10m > /var/tmp/k-sched-nice.log 2>&1 &
$ renice -n 10 -p $(pgrep -nx stress-ng-cpu)
252038 (process ID) old priority 0, new priority 10
$ pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:03 UID PID %usr %system %guest %wait %CPU CPU Command 09:35:08 1001 252037 90.24 0.00 0.00 9.56 90.24 1 stress-ng-cpu 09:35:08 1001 252038 9.76 0.00 0.00 89.84 9.76 1 stress-ng-cpu …
$ grep se.load.weight /proc/$(pgrep -ox stress-ng-cpu)/sched /proc/$(pgrep -nx stress-ng-cpu)/sched
/proc/252037/sched:se.load.weight : 1048576 /proc/252038/sched:se.load.weight : 112640

pgrep -n picks the newest worker, and renice raised its nice value to 10. The measured split is 90.24% to 9.76%, and each worker's %wait is the other's share: at any moment one runs and one waits. se.load.weight reads 112640 for the reniced worker, 110 times 1024. Nice changes a task's share, not its right to run: if the nice 0 worker stopped, the other would get the whole CPU. You may raise the nice value of your own processes; lowering it again, even back to 0, needs root (CAP_SYS_NICE) or an RLIMIT_NICE allowance.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop k-sched-nice.scope

Between services, cgroup weights decide

When the cgroup cpu controller is enabled, the scheduler works in levels. It first divides CPU time between sibling cgroups according to their cpu.weight (default 100, range 1 to 10000), then between the tasks inside each group according to nice. Here two services each run one worker on CPU 1, and the batch worker is reniced to 19.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=k-sched-web --uid=$USER stress-ng --cpu 1 --taskset 1 --timeout 10m > /var/tmp/k-sched-web.log 2>&1 & sudo systemd-run --scope --unit=k-sched-batch --uid=$USER stress-ng --cpu 1 --taskset 1 --timeout 10m > /var/tmp/k-sched-batch.log 2>&1 &
$ renice -n 19 -p $(pgrep --cgroup /system.slice/k-sched-batch.scope -x stress-ng-cpu)
252196 (process ID) old priority 0, new priority 19
$ pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:11 UID PID %usr %system %guest %wait %CPU CPU Command 09:35:16 1001 252195 50.20 0.00 0.00 50.00 50.20 1 stress-ng-cpu 09:35:16 1001 252196 50.00 0.00 0.00 50.00 50.00 1 stress-ng-cpu …

The renice changed nothing: 50.20% and 50.00%. Each scope is its own group with weight 100, and nice only ranks tasks inside the batch group, which holds one task. Whether services are separate groups depends on whether systemd has enabled the cpu controller, the part of the cgroup code that divides CPU time, in system.slice. It does that for every unit in the slice as soon as one of them sets a CPU property. A cgroup's cgroup.subtree_control file lists the controllers it passes down to its children:

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/fs/cgroup/system.slice/cgroup.subtree_control grep CPUWeight /usr/lib/systemd/system/multipathd.service
cpu memory pids CPUWeight=1000

On a default Ubuntu Server install, the enabled multipathd.service sets CPUWeight=1000, so the controller is on and nice never crosses a service boundary; a host where multipathd is disabled may behave like Rocky below. The Rocky Linux 10 lab machine has no unit like that: system.slice passes no cpu controller down, all services share one group, and the same experiment with nice 10 splits 90 to 10. Check cgroup.subtree_control on your own hosts before relying on nice between services.

deploy@rocky10 · Rocky Linux 10.2
$ cat /sys/fs/cgroup/system.slice/cgroup.subtree_control
memory pids
$ pidstat -u -C stress-ng-cpu 5 1
Linux 6.12.0-211.16.1.el10_2.0.1.aarch64 (rocky10) 09/27/2026 _aarch64_ (2 CPU) 09:26:35 AM UID PID %usr %system %guest %wait %CPU CPU Command 09:26:40 AM 1001 54221 9.74 0.00 0.00 89.46 9.74 1 stress-ng-cpu 09:26:40 AM 1001 54222 90.26 0.00 0.00 9.74 90.26 1 stress-ng-cpu …

Between services, the setting to use is CPUWeight=, which systemd writes to cpu.weight.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl set-property --runtime k-sched-batch.scope CPUWeight=20 cat /sys/fs/cgroup/system.slice/k-sched-web.scope/cpu.weight /sys/fs/cgroup/system.slice/k-sched-batch.scope/cpu.weight
100 20
$ pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:19 UID PID %usr %system %guest %wait %CPU CPU Command 09:35:24 1001 252195 83.63 0.00 0.00 16.77 83.63 1 stress-ng-cpu 09:35:24 1001 252196 16.57 0.00 0.00 83.43 16.57 1 stress-ng-cpu …
$ sudo systemctl stop k-sched-web.scope pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:24 UID PID %usr %system %guest %wait %CPU CPU Command 09:35:29 1001 252196 100.40 0.00 0.00 0.00 100.40 1 stress-ng-cpu …

Weights of 100 and 20 predict 83.3% and 16.7%; pidstat measured 83.63% and 16.57%. A weight is work-conserving: it only matters while groups compete, so when the web scope stopped, the batch worker at weight 20 took all of CPU 1. A quota is not work-conserving.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl set-property --runtime k-sched-batch.scope CPUQuota=25% cat /sys/fs/cgroup/system.slice/k-sched-batch.scope/cpu.max
25000 100000
$ mpstat -P 1 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:31 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle 09:35:36 1 0.20 25.25 0.00 0.00 0.00 0.00 0.00 0.00 0.00 74.55 …

CPUQuota=25% became cpu.max 25000 100000: 25 ms of CPU time per 100 ms period, after which the group is throttled until the next period even when the CPU has nothing else to do. CPU 1 is now about 75% idle while the batch job waits. Its time appears as %nice because the worker still runs at nice 19. Use weights to protect one service from another and a quota when a hard ceiling is the requirement; the resource-control lesson covers the systemd settings in depth.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop k-sched-batch.scope

Real-time policies

A SCHED_FIFO task, priority 1 to 99, runs until it blocks, yields or a higher-priority real-time task becomes runnable. Fair tasks get only the time real-time tasks leave, apart from one reservation shown below. SCHED_RR is the same but rotates tasks of equal priority every kernel.sched_rr_timeslice_ms (100 ms). Setting either needs CAP_SYS_NICE, so sudo, unless RLIMIT_RTPRIO allows it. Here a fair worker and a worker switched to SCHED_FIFO share CPU 1.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=k-sched-web --uid=$USER stress-ng --cpu 1 --taskset 1 --timeout 10m > /var/tmp/k-sched-web.log 2>&1 & sudo systemd-run --scope --unit=k-sched-rt --uid=$USER stress-ng --cpu 1 --taskset 1 --timeout 10m > /var/tmp/k-sched-rt.log 2>&1 &
$ rt=$(pgrep --cgroup /system.slice/k-sched-rt.scope -x stress-ng-cpu) sudo chrt --fifo --pid 10 $rt chrt -p $rt
pid 252492's current scheduling policy: SCHED_FIFO pid 252492's current scheduling priority: 10
$ pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:38 UID PID %usr %system %guest %wait %CPU CPU Command 09:35:44 1001 252492 95.41 0.00 0.00 5.19 95.41 1 stress-ng-cpu 09:35:44 1001 252493 4.99 0.00 0.00 94.81 4.99 1 stress-ng-cpu …

The FIFO worker took 95.41% and the fair worker kept 4.99%. Older documentation credits that 5% to RT throttling through kernel.sched_rt_runtime_us, 950000 of every 1000000 microseconds. Since kernel 6.12, on kernels built without RT group scheduling (Ubuntu and RHEL both leave it off), that sysctl only limits admission of SCHED_DEADLINE tasks. The 5% now comes from the fair server: a deadline-class reservation on each CPU that runs waiting fair tasks for 50 ms of every second (its runtime and period files in debugfs are in nanoseconds).

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/sys/kernel/sched_rt_runtime_us sudo cat /sys/kernel/debug/sched/fair_server/cpu1/runtime /sys/kernel/debug/sched/fair_server/cpu1/period
950000 50000000 1000000000
$ sudo systemctl stop k-sched-web.scope pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:35:44 UID PID %usr %system %guest %wait %CPU CPU Command 09:35:49 1001 252492 100.00 0.00 0.00 0.00 100.00 1 stress-ng-cpu …

With no fair task waiting, the FIFO worker got the whole CPU: the reservation only takes time when fair tasks need it. RHEL 10's 6.12 kernel measured the same pattern: 94.40% and 6.00%, then 100.20% alone. Stop the scope with sudo systemctl stop k-sched-rt.scope.

Real-time is for short, bounded work
Give SCHED_FIFO or SCHED_RR only to work that runs briefly and then sleeps, such as an audio or control loop. A CPU-bound real-time task on every CPU leaves everything else, your SSH session included, 50 ms per second on each CPU. The kernel documentation describes that reservation as a little time to recover the machine from a runaway real-time task, not time to run services.

Measuring run-queue latency

Run-queue latency is the time a task spends runnable but not running: from the moment it wakes up or is preempted until it gets a CPU. The next setup crowds CPU 1 once more with two CPU-bound workers and a probe, stress-ng --cyclic, which sleeps 1 ms in a loop like a service handling small requests (--cyclic-sleep is in nanoseconds). The kernel truncates process names to 15 characters, so the probe shows up as stress-ng-cycli.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=k-sched-batch --uid=$USER stress-ng --cpu 2 --taskset 1 --timeout 10m > /var/tmp/k-sched-batch.log 2>&1 & sudo systemd-run --scope --unit=k-sched-probe --uid=$USER stress-ng --cyclic 1 --cyclic-policy other --cyclic-sleep 1000000 --taskset 1 --timeout 10m > /var/tmp/k-sched-probe.log 2>&1 &
$ pidstat -w -C stress-ng 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:36:00 UID PID cswch/s nvcswch/s Command 09:36:05 1001 252681 941.72 0.00 stress-ng-cycli 09:36:05 1001 252682 0.00 612.38 stress-ng-cpu 09:36:05 1001 252683 0.00 612.18 stress-ng-cpu …

pidstat -w counts context switches per second. Each CPU-bound worker is switched out about 610 times a second without asking (nvcswch/s): it is preempted. The probe gives up the CPU about 940 times a second voluntarily (cswch/s), once for every sleep. Mostly nonvoluntary switches mean a task is waiting for a CPU; mostly voluntary ones mean it waits for something else.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/$(pgrep -ox stress-ng-cpu)/schedstat /proc/$(pgrep -nx stress-ng-cycli)/schedstat
7587473489 7668420453 9155 75051323 71484924 14305

The second field of schedstat is the total time in nanoseconds a task waited in the run queue, and the third is how many times it was scheduled. The first worker waited 7.67 s over 9155 turns, about 0.84 ms each. The probe waited 71.5 ms over 14305 wakeups, about 5 microseconds each. That is EEVDF's lag at work: the probe uses a tiny part of its share, so whenever it wakes it is eligible and has the earliest deadline. A histogram shows the whole distribution. On Ubuntu 26.04 the bcc version of the tool fails:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo timeout -s INT 5 runqlat-bpfcc
… include/linux/fs.h:2431:15: error: static assertion failed due to requirement 'sizeof(struct filename) % 64 == 0': sizeof(struct filename) % 64 == 0 … Exception: Failed to compile BPF module <text>

bcc compiles each tool against the kernel headers when it starts, and the bcc 0.35 in the Ubuntu archive fails to compile this tool against the kernel 7.0 headers. bpftrace, installed by default on Ubuntu Server, ships the same tool as runqlat.bt; it reads kernel types from the kernel's built-in BTF data instead of headers. On RHEL 10 the bcc version works as /usr/share/bcc/tools/runqlat. Both run until you press Ctrl-C; timeout -s INT 5 does that after five seconds. The tool runs on every wakeup and context switch on every CPU, and its man page warns that for some workloads its overhead becomes significant: measure it on a test machine first and keep production runs short.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo timeout -s INT 5 runqlat.bt
Attached 5 probes Tracing CPU scheduler... Hit Ctrl-C to end. @usecs: [0] 9 | | [1] 2449 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ | [2, 4) 1330 |@@@@@@@@@@@@@@@@@@@@@@@ | [4, 8) 755 |@@@@@@@@@@@@@ | [8, 16) 434 |@@@@@@@ | [16, 32) 142 |@@ | [32, 64) 47 | | [64, 128) 18 | | [128, 256) 28 | | [256, 512) 349 |@@@@@@ | [512, 1K) 1265 |@@@@@@@@@@@@@@@@@@@@@@ | [1K, 2K) 2975 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| [2K, 4K) 12 | |

Read the histogram as two populations. The short waits, 1 to 16 microseconds, are the probe and the system's own tasks. A second peak between 0.5 and 2 ms is the two workers waiting for each other's turn, about one base slice of 1.4 ms. An average would have hidden that shape.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo perf sched record -o /var/tmp/k-sched.perf -- sleep 3 sudo perf sched latency -i /var/tmp/k-sched.perf | grep -E "Task|stress-ng"
[ perf record: Woken up 1 times to write data ] [ perf record: Captured and wrote 2.104 MB /var/tmp/k-sched.perf (18111 samples) ] Task | Runtime ms | Count | Avg delay ms | Max delay ms | Max delay start | Max delay end | stress-ng-cpu:(2) | 2983.268 ms | 3360 | avg: 0.903 ms | max: 16.730 ms | max start: 4210.221260 s | max end: 4210.237990 s stress-ng-cycli:252681 | 26.396 ms | 2676 | avg: 0.017 ms | max: 7.320 ms | max start: 4210.271456 s | max end: 4210.278777 s

perf sched records every scheduler event and perf sched latency summarises them per task. It needs sudo; on Ubuntu, kernel.perf_event_paranoid is 4, which blocks every unprivileged use of perf. The workers waited 0.903 ms on average; the probe waited 0.017 ms on average. The maxima, 16.7 ms and 7.3 ms, are single events in three seconds: an average hides them, and a histogram shows how rare they are. The recording's size grows with the number of CPUs times the switch rate: 2.1 MB for three seconds on this two-vCPU machine, many times that on a large busy host. There, record for one to three seconds with -o on a filesystem with room, and check the output for lost events.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/fs/cgroup/system.slice/k-sched-batch.scope/cpu.pressure
some avg10=93.75 avg60=37.15 avg300=9.14 total=28667574 full avg10=0.18 avg60=0.03 avg300=0.00 total=215891

The scope's cpu.pressure gives the same picture as a share of time: in the last ten seconds, at least one of its tasks was waiting for a CPU 94% of the time (some). full, all of its tasks waiting at once, stays near zero because one of the two workers is almost always running. The next step follows from which tasks wait. If they are the ones users wait for, separate them from the competing work with affinity, a cpuset or weights, or reduce that work. If a slow service shows microsecond run-queue delays, the scheduler is not the cause, and the next place to look is locks, I/O or whatever the service calls.

Try this

With the batch and probe scopes still running, move one worker to CPU 0 with taskset -cp 0 $(pgrep -nx stress-ng-cpu) and run pidstat -u -w -C stress-ng-cpu 5 1. Predict first. Expect both workers at about 100%, each on its own CPU, with %wait of a percent or two. The worker still on CPU 1 keeps about 960 nonvoluntary switches a second, because the probe wakes a thousand times a second and preempts it; the moved worker shows fewer than 20. Then stop both scopes with sudo systemctl stop k-sched-batch.scope k-sched-probe.scope and delete the recording with sudo rm /var/tmp/k-sched.perf.

Takeaway

Measure the wait (schedstat, runqlat, CPU pressure) before you change priorities, then choose the lever that matches the boundary: nice inside a service, CPUWeight= between services, a quota for a hard ceiling, affinity to separate work, and real-time policies only for short, bounded work.

Quick check
01Users report slow API responses while a batch job runs on the same host. perf sched latency shows the API worker threads with an average delay of 0.02 ms and a maximum of 0.3 ms, while the batch threads average 0.9 ms. What do you conclude?
Incorrect — Starved threads would show long run-queue delays. Delays of microseconds mean the API threads get a CPU almost as soon as they want one.
Correct — The scheduler hands the API threads a CPU within microseconds, so the time is going somewhere else: locks, I/O or the services the API calls.
Incorrect — The batch delays are expected for CPU-bound work sharing CPUs, and giving the API more weight cannot fix a delay the API does not have.
Incorrect — perf sched records per-task events, and threads are tasks. Its figures agree with runqlat, which only adds the distribution.
02On Ubuntu 26.04 you run a CPU-bound test program as SCHED_FIFO on CPU 1, with nothing else runnable there. mpstat shows CPU 1 at 100%. A colleague expected 95% because kernel.sched_rt_runtime_us is 950000. Who is right, and why?
Incorrect — pidstat and mpstat both measured the whole CPU going to the FIFO task. The throttling model the colleague uses no longer applies to these kernels.
Incorrect — sched_rt_runtime_us never distinguished FIFO from RR, and on these kernels it no longer throttles either of them.
Incorrect — That describes borrowing between CPUs, which is not what happens here. Without RT group scheduling the sysctl does not throttle real-time tasks at all.
Correct — Since 6.12 the fair server protects fair tasks with 50 ms per second on each CPU, and it only takes that time when fair tasks are waiting.
03A batch service should use spare CPU at night but yield to the web service during the day, with no schedule changes. Both run under system.slice on Ubuntu Server. Which setting fits?
Correct — Weights only act while groups compete, so at night the batch job gets whatever is free and by day the web service wins the split.
Incorrect — A quota caps the batch job at night too, leaving the CPUs idle while it waits. It is a ceiling, not a priority.
Incorrect — On a default Ubuntu Server install, where multipathd enables the cpu controller, every service is its own cpu cgroup, so nice only ranks the batch job's own tasks.
Incorrect — A real-time web service would preempt everything, including the system's own fair tasks, whenever it is busy. That is not a priority knob for services.

Related