Debugging Pending pods: the five usual causes

Insufficient resources, taints, affinity, PVCs, and quotas — how to read the events and fix each one fast.

Jul 16, 2024·Updated ·4 min readBeginner·By SecOpsLog · command-tested

Pending means the scheduler has not placed the Pod on a node, and the scheduler is unusually forthcoming about why: it writes the reason into a FailedScheduling event on the Pod, with the count of nodes it considered and the filter each one failed. The sequence below reads that event first and treats every other command as confirmation. Deleting and recreating the Pod is the step to skip; a Pod that cannot fit comes back with the same event and a new name.

bash — representative: one event on a six-node cluster (two filters at once)
kubectl describe pod api-7c9f2 -n prod | grep -A3 "^Events"
Events:
Type Reason Age From Message
Warning FailedScheduling 41s default-scheduler 0/6 nodes are available: 2 node(s) had untolerated taint(s), 4 Insufficient cpu. preemption: 0/6 nodes are available: 4 No preemption victims found for incoming pod, 2 Preemption is not helpful for scheduling.
six nodes, two rejected by a taint, four by CPU: the fix is capacity or requests, not the taint
bash — observed: each reason on its own, one node, Kubernetes 1.35.0 (the message the scheduler actually writes)observed
# a Pod requesting 100 CPU cores
0/1 nodes are available: 1 Insufficient cpu. ... preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.
# the node tainted dedicated=gpu:NoSchedule, a Pod with no toleration
0/1 nodes are available: 1 node(s) had untolerated taint(s). ...
# a nodeSelector no node carries
0/1 nodes are available: 1 node(s) didn't match Pod's node affinity/selector. ...
# a PVC that names a StorageClass which does not exist
0/1 nodes are available: pod has unbound immediate PersistentVolumeClaims. ...
# a Pod with schedulingGates
phase=Pending PodScheduled.reason=SchedulingGated FailedScheduling event: <none>
on 1.35 the taint line no longer prints {dedicated: gpu}; it says "untolerated taint(s)". The specific taint is in `kubectl describe node`, not the scheduler summary
Two places placement stops: admission and scheduling

A ResourceQuota refuses the Pod at admission, before it exists, so there is no Pending Pod and the FailedCreate lands on the ReplicaSet. A Pod that is admitted either waits on a schedulingGate with no event, or reaches the scheduler, where each filter removes nodes and the FailedScheduling event names the one that emptied the list.

Kubernetes Pending triage: kubectl apply first passes admission, where a ResourceQuota can refuse the Pod with a 403 so it is never created and the FailedCreate message lands on the ReplicaSet rather than on a Pending Pod; an admitted Pod with schedulingGates shows SchedulingGated with no event until the gate is cleared; otherwise the scheduler runs its filters (insufficient cpu or memory, untolerated taints, node affinity or selector mismatch, unbound PersistentVolumeClaims) and either a node fits and the Pod runs, or none fits and the Pod stays Pending with a FailedScheduling event naming the filter kubectl apply -f deploy.yamlAdmission: ResourceQuotaschedulingGates present?Scheduler filtersInsufficient cpu / memoryuntolerated taint(s)node affinity / selectorunbound PVCadmitted: Pod creatednone setOver quota: 403 at admissionobject never created — no PodReplicaSet FailedCreate; read the RSover quotaSchedulingGated: no eventthe Pod exists, is not scheduledthe gate owner must clear itgates seta node fits: Scheduled, RunningPendingFailedScheduling names the filterTwo different failures: admission refuses before a Pod exists (no Pending, read the ReplicaSet);scheduling leaves a Pending Pod whose FailedScheduling event names the filter to fix.

The messages, and the fix for each

FailedScheduling messages mapped to causes

Message fragmentCauseConfirm withFix
Insufficient cpu / Insufficient memoryno node has enough allocatable left after existing requestskubectl describe node (Allocatable and Allocated resources)lower the requests, add nodes, or let the Cluster Autoscaler
node(s) had untolerated taint(s)the nodes with room are reserved for something else (the summary no longer names the taint on 1.35)kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taintsa matching toleration, if the Pod belongs on those nodes
node(s) didn't match Pod's node affinity/selectorthe selector names a label no schedulable node carrieskubectl get nodes --show-labelsfix the label or the selector; watch for typos and case
pod has unbound immediate PersistentVolumeClaimsa PVC the Pod mounts is not Boundkubectl describe pvcfix the StorageClass, provisioner or topology; the Pod follows
node(s) didn't match pod anti-affinity rules / topology spread constraintsthe rule leaves no valid node given where the other replicas arekubectl get pods -o wide for the same apprelax the rule or add a node in the missing zone
status SchedulingGated, no eventthe Pod carries schedulingGates and something is expected to remove themkubectl get pod -o jsonpath='{.spec.schedulingGates}'the controller that owns the gate, or remove it if it is stale

Capacity is the message people misread. The scheduler compares requests against a node's allocatable, which is capacity minus what the kubelet reserves for system daemons and eviction thresholds, and then minus the requests of every Pod already on the node. A node showing 40 % CPU usage in monitoring can still be full from the scheduler's point of view, because scheduling is about requests, not usage. kubectl describe node prints both numbers under Allocatable and Allocated resources, and the difference between them is what a new Pod can ask for.

bash — representative: capacity, allocatable, and what is already promised (a busy worker)
kubectl describe node worker-3 | grep -A6 "^Allocatable" | grep -E "cpu|memory"
cpu: 7500m
memory: 30Gi
kubectl describe node worker-3 | grep -A4 "Allocated resources" | grep -E "cpu|memory"
cpu 7200m (96%) 3400m (45%)
memory 22Gi (73%) 28Gi (93%)
requests at 96 %: a 500m Pod does not fit here although the node is idle; the second column is limits

The case that is not Pending: quota

A namespace ResourceQuota is enforced at admission, before the scheduler sees anything. A Pod that would push the namespace over its quota is rejected with a 403 and is never created, so there is no Pending Pod and no FailedScheduling event to read. The symptom is a Deployment whose replica count never reaches its target, and the message is on the ReplicaSet, which is the object that tried to create the Pod and was refused. kubectl events on the namespace shows it without knowing which object to describe.

bash — observed: quota rejection lives on the ReplicaSet, not on a Pending Pod (quota requests.cpu=1, three 500m replicas)observed
kubectl get deploy api -n shop -o jsonpath="{.status.availableReplicas}/{.spec.replicas}"
2/3
kubectl get event -n shop --field-selector reason=FailedCreate -o jsonpath="{.items[-1:].message}"
(combined from similar events): Error creating: pods "api-5c49ccb4c7-8s8cf" is forbidden: exceeded quota: shop-quota, requested: requests.cpu=500m, used: requests.cpu=1, limited: requests.cpu=1
kubectl get pod -n shop --field-selector status.phase=Pending -o name | wc -l
0
two of three Pods fit under the 1-core quota; the third is refused at admission and never created — no Pending Pod, no FailedScheduling event. Raise the quota or reduce requests elsewhere

The order that avoids the wrong fix

Read the Pod's events. If there are none and the status is SchedulingGated, find the gate's owner. If there are none and the Pod does not exist at all, read the ReplicaSet or the Job. If the event names a filter, confirm it with the one command in the table and change the thing the filter checks. Adding a toleration to get past a taint is only correct when the Pod belongs on the tainted nodes; lowering a request to get past Insufficient memory is only correct when the request was inflated. The event tells you which filter said no; whether the filter was right is the engineering question.

Preemption changes who is Pending
When the Pending Pod has a higher PriorityClass than Pods on a full node, the scheduler may evict those to make room, and the message ends with the preemption outcome. A Pod that suddenly went Pending after being Running was often preempted by something more important; its own event says so, and the fix is priority design rather than capacity.
What was run for this article
A disposable single-node kind v0.31.0 cluster (Kubernetes 1.35.0, node image by digest; Docker Engine 28.5.2, linux/arm64; build/evidence/k8s-troubleshoot-pending). Each case is created on its own so the scheduler names one reason: a Pod requesting 100 CPU cores (Insufficient cpu), the node tainted with a Pod that has no toleration (untolerated taint(s)), a nodeSelector no node carries (node affinity/selector), a PVC on a missing StorageClass (unbound PersistentVolumeClaims), a schedulingGates Pod (SchedulingGated, no event), and a Deployment over a 1-core ResourceQuota (a 403 FailedCreate on the ReplicaSet, no Pending Pod). Seven assertions check each message fragment; the exact wording is 1.35.0's. The six-node event, the worker-3 capacity numbers, the anti-affinity/topology row and the preemption note are representative: one kind node cannot show them.

Pending ends when the Pod is placed; the next two statuses have their own triage. CrashLoopBackOff is the process failing after placement, which is usually a probe or startup problem or a limit set below what the process needs; ImagePullBackOff is the registry, the tag or the pull secret, and the event names that too.

Related posts

Quick reference