The cluster: control plane & nodes
The two halves of every cluster.
A Kubernetes cluster is a group of computers, real or virtual, tied together so you can treat them as one system. It splits into two jobs: the control plane decides and records desired state; the worker nodes run the workloads. Neither side does the other's job. They stay in sync through the API.
Two halves, two jobs
Every cluster splits into two roles. The control plane decides and remembers. Worker nodes run your programs. A node is one machine; a worker's job is running apps. A practice cluster on a laptop might put both roles on one machine. Production can spread them across dozens or hundreds. The split is about jobs, not machine count.
Your apps run as containers. A container is a program packed up with everything it needs, so it behaves the same on your laptop and on a server in a data centre. Kubernetes never hands you a bare container, though. It wraps one or more containers in a Pod, the smallest unit Kubernetes will schedule and start. Most Pods hold a single container, so for now you can read one Pod as one running program: a small box with your container inside it.
Here is where almost every beginner trips. You do not log into the worker machines to start your apps. You talk to the control plane and nothing else, and it arranges for the workers to do the work. Want three copies of your app running at all times? You say that to the control plane, and it works out where and how. The tool you use for that conversation is kubectl (say it "kube control"), the command-line program that carries your requests to the cluster. It gets a lesson of its own next.
kubectl get nodes
NAME STATUS ROLES AGE VERSIONcp-1 Ready control-plane 9d v1.31.0worker-1 Ready <none> 9d v1.31.0worker-2 Ready <none> 9d v1.31.0
Read that output top to bottom. cp-1 has the control-plane role. The two workers show <none> under ROLES: ordinary workers whose job is running apps. Ready means the machine is healthy and available. AGE is how long it has been in the cluster. VERSION is the Kubernetes version it runs. You talk to cp-1. It tells worker-1 and worker-2 what to run.
What the control plane is made of
You do not need to memorize internals yet, but a short tour helps. The control plane is a handful of programs. The API server is the only entry point; every request goes through it, including yours. (API stands for Application Programming Interface: the doorway one program uses to talk to another.) Behind it, etcd stores the record of what is supposed to exist. The scheduler picks which worker node each new Pod lands on. Controllers keep comparing what is running to what you asked for and closing the gap. A Pod dies at 3am, a controller sees it, a replacement starts. Nobody gets paged.
kubectl get pods -n kube-system
NAME READY STATUS RESTARTS AGEcoredns-7db6d8ff4d-abcde 1/1 Running 0 9detcd-cp-1 1/1 Running 0 9dkube-apiserver-cp-1 1/1 Running 0 9dkube-controller-manager-cp-1 1/1 Running 0 9dkube-proxy-2xk9p 1/1 Running 0 9dkube-scheduler-cp-1 1/1 Running 0 9d
The -n kube-system flag points kubectl at kube-system, the namespace where Kubernetes keeps its own Pods. Control-plane pieces run there as Pods: kube-apiserver, etcd, kube-scheduler and kube-controller-manager, plus coredns (cluster DNS, Domain Name System: names to addresses) and kube-proxy for networking. On a managed cloud cluster, the control-plane entries are often missing from this list because the provider runs the control plane for you.
What a worker node does all day
Each worker runs a kubelet. That agent takes instructions from the control plane, starts the containers it is told to start, watches them, and reports health back. A networking component on the node wires Pods so they can reach each other and the outside world. Workers wait for assignments and carry them out.
That steady reporting doubles as the cluster's early-warning system. When a worker reboots, fills its disk, or falls off the network, its kubelet goes quiet, and the control plane notices the silence. Run the same command on a bad day and the machine tells on itself:
kubectl get nodes
NAME STATUS ROLES AGE VERSIONcp-1 Ready control-plane 9d v1.31.0worker-1 Ready <none> 9d v1.31.0worker-2 NotReady <none> 9d v1.31.0
worker-2 has gone NotReady. To find out why, ask that node to describe itself. The Conditions and Events near the bottom of the output spell it out in plain words:
kubectl describe node worker-2
Name: worker-2Conditions:Type Status Reason Message---- ------ ------ -------Ready Unknown NodeStatusUnknown Kubelet stopped posting node status.Events:Type Reason Age From Message---- ------ ---- ---- -------Normal NodeNotReady 3m node-controller Node worker-2 status is now: NodeNotReady
Status Unknown, with "Kubelet stopped posting node status", means the machine went quiet and the control plane no longer trusts it. The event comes from node-controller, acting without a human. After a short grace period it moves worker-2's Pods onto healthy nodes, so the app keeps serving while you check whether that box rebooted or filled its disk.
kubectl cluster-info
Kubernetes control plane is running at https://192.168.49.2:8443CoreDNS is running at https://192.168.49.2:8443/api/v1/namespaces/kube-system/services/kube-dns:dns/proxyTo further debug and diagnose cluster problems, use 'kubectl cluster-info dump'.
That address is the API server, the front door you send everything to. Notice there is one URL for the whole cluster, not one per machine. You never point kubectl at a worker. Send your request to that single address and the control plane takes over: it writes down what you asked for, the scheduler picks a node, and the kubelet on that node starts your containers. You describe the goal. The cluster works out the steps.
Every cluster has two jobs: decide, and do. The control plane (API server, etcd, scheduler, controller manager) decides. The worker nodes run the kubelet, a container runtime (the software that actually starts and stops containers) and usually a network agent, and they do. When something breaks, the first useful question is which half owns the symptom. Answer that and you stop firing off random kubectl commands and hoping.
etcd is the durable memory of the cluster, but only the API server is allowed to talk to it. Everything else goes through the API: kubectl, the controllers, the kubelets on every node. That is why a sick API server feels like the whole world is on fire even when your apps are still serving. It is also why protecting and backing up that one path is worth far more of your attention than any pretty dashboard.
Say your production cluster shows every node Ready but Pods sitting in Pending. That usually points at the scheduler or at plain capacity, not at "Kubernetes is down". Read the node conditions and the events first. Rebooting control-plane virtual machines in a panic is how a small problem turns into a real outage.
Try this
Map your own cluster. List the nodes with their roles and conditions, then list the control-plane Pods in kube-system, and you can see exactly who lives where.
$ kubectl get nodes -o wideNAME STATUS ROLES AGE VERSION INTERNAL-IPkind-control-plane Ready control-plane 12d v1.30.0 172.18.0.2kind-worker Ready <none> 12d v1.30.0 172.18.0.3$ kubectl get pods -n kube-system -o wide | head -20NAME READY STATUS NODEcoredns-... 1/1 Running kind-control-planeetcd-kind-control-plane 1/1 Running kind-control-planekube-apiserver-kind-control-plane 1/1 Running kind-control-planekube-controller-manager-kind-control-plane 1/1 Running kind-control-planekube-scheduler-kind-control-plane 1/1 Running kind-control-planekube-proxy-... 1/1 Running kind-worker$ kubectl describe node kind-worker | Select-String -Pattern 'Conditions:|Ready|MemoryPressure|DiskPressure' -Context 0,6Ready True
Takeaway
Control plane decides, nodes execute. When a symptom shows up, name the half that owns it before you touch anything. A Pod stuck in Pending is a scheduling question. A node stuck at NotReady is a kubelet question. A kubectl command that hangs is an API server question. Aim at that one layer instead of rebooting machines and hoping the symptom goes away.