It is 3 a.m. and a node in your production cluster runs low on memory. When the dust settles, your payment service has been killed and restarted twice. The nightly batch job on the same node, the one nobody would miss for an hour, sailed through untouched. Nothing in the dashboards says why, and nobody made that choice on purpose.
Except somebody did. Months earlier, someone wrote requests and limits into two manifests. The batch job got requests equal to its limits. The payment service got a small memory request and a generous limit. Those few lines quietly told Kubernetes which workload to sacrifice first, and Kubernetes followed the instructions exactly.
The mechanism behind that decision is the Quality of Service (QoS) class. This post follows it from the fields you write in YAML down to the kernel settings that actually decide who lives. Version notes are checked against the Kubernetes v1.37 documentation.
Why Kubernetes Needs a Pecking Order
The scheduler places pods by their requests, not their actual usage. A node with 16 GiB of allocatable memory can host pods whose requests add up to 16 GiB, while their limits add up to 40 GiB. This is overcommitment, and it is deliberate. Most workloads use far less than their limits most of the time, and packing nodes tightly is how clusters stay affordable.
The catch is that overcommitment is a bet. When enough pods burst at once, the node runs out of something real: CPU time, memory, disk, or process IDs. CPU is forgiving because it can be shared and slowed down. Memory is not. You cannot slow down an allocation; eventually something has to die.
So every node needs a policy for who suffers first. Kubernetes could have asked you to declare that policy directly. Instead, it infers it from the resource fields you already write. The QoS class is that inference.
The Model: Three Classes You Never Set
There is no qosClass field in a pod spec. The API server computes the class when the pod is created and writes it to status.qosClass. You influence it only through resources, and the rules are strict.
- Guaranteed: every container in the pod has CPU and memory requests and limits, all greater than zero, with each request equal to its limit.
- Burstable: the pod is not Guaranteed, but at least one container has at least one CPU or memory request or limit.
- BestEffort: no container sets any CPU or memory request or limit at all.
A few details in those rules catch people out. Init containers count, including sidecars, which are implemented as init containers. One sidecar with no limits drops an otherwise perfect pod to Burstable. Only CPU and memory matter; a pod can request GPUs or ephemeral storage and still be BestEffort.
Defaulting also plays a part. If you set a limit without a request, Kubernetes copies the limit into the request. A container with only limits for both CPU and memory is therefore Guaranteed, which surprises people who thought they were being cautious. A namespace LimitRange can inject defaults too, so the class of a pod can depend on where it is deployed.
# Guaranteed: request equals limit for both resources
resources:
requests: { cpu: "1", memory: "2Gi" }
limits: { cpu: "1", memory: "2Gi" }
# Burstable: requests below limits
resources:
requests: { cpu: "250m", memory: "512Mi" }
limits: { cpu: "2", memory: "2Gi" }
# Also Guaranteed: requests default to the limits
resources:
limits: { cpu: "500m", memory: "1Gi" }
The third block is the trap. It looks like "only set a ceiling," but the defaulting rule turns it into a full reservation. The scheduler will reserve 500m of CPU and 1 GiB of memory for that container, whether it uses them or not.
Inside the Node: The cgroup Tree
The class only matters because the kubelet acts on it. When a pod lands on a node, the kubelet builds a cgroup for it, and where that cgroup sits in the hierarchy depends on the QoS class. On a systemd-managed cgroup v2 node, the layout looks like this.
The per-class parent slices exist so the kubelet can manage each class as a group. For example, it can keep the BestEffort slice's CPU weight at the floor no matter how many BestEffort pods are running. On older cgroup v1 nodes, the paths look different (/sys/fs/cgroup/memory/kubepods/burstable/...), but the structure is the same.
Everything that follows comes back to the numbers in the bottom row. Each one is a different mechanism, and each responds to a different kind of pressure.
CPU: Requests Become Weight, Limits Become a Ceiling
CPU requests and CPU limits turn into two unrelated kernel controls, and most confusion about CPU in Kubernetes comes from treating them as the same thing.
A CPU request becomes a weight. The kubelet converts millicores into CFS shares (request_millicores × 1024 / 1000, with a minimum of 2), and the container runtime maps those onto cpu.weight on cgroup v2. Weight only matters under contention. If two containers want the CPU at the same moment, they split it in proportion to their weights. If the CPU is idle, a container with a tiny weight can still use every core. BestEffort pods get the minimum weight, so they receive leftovers only.
A CPU limit becomes a quota, written to cpu.max as "this many microseconds of CPU time per period." The default period is 100 ms, so a 500m limit becomes 50000 100000: 50 ms of CPU time in every 100 ms window. Once the container spends its quota, the kernel stops scheduling it until the next period begins, even if every other core on the machine is idle. This is CFS throttling.
Quota is measured in CPU time across all threads, so parallelism burns through it faster. A service with 4 busy threads and a 500m limit spends its entire 50 ms allowance in about 12.5 ms of wall-clock time, then sits frozen for the remaining 87.5 ms. Averaged over a minute, the container looks like it uses exactly half a core. Measured per request, some requests take 90 ms longer than they should.
This is why QoS and CPU interact in a way that seems backwards. Guaranteed pods must have a CPU limit equal to their request. They get strong memory protection, but they are also the pods most likely to be throttled, because they have no headroom above their request at all. Burstable pods without a CPU limit can soak up idle cycles freely and are never throttled by quota.
Memory: The Limit Is a Wall, the Request Is a Promise
Memory works differently because memory cannot be throttled into submission the way CPU can. A memory limit becomes memory.max. When a container tries to allocate past it, the kernel first tries to reclaim page cache inside that cgroup. If that fails, it invokes the OOM killer scoped to that cgroup. This is the familiar OOMKilled status, and it has nothing to do with QoS class or with what other pods are doing. A Guaranteed pod that exceeds its own limit dies just like any other.
A memory request, by default, is not enforced by the kernel at all. It affects scheduling and it affects the two triage mechanisms described next. That is what makes the request a promise rather than a reservation: the node promises to try to keep you at or above it, and it uses your class to decide how hard it tries.
There is one newer exception. With Memory QoS on cgroup v2 (beta since v1.37, and off until configured), the kubelet can translate requests into kernel-level protection. It sets memory.min for Guaranteed pods and memory.low for Burstable pods, and it can use memory.high to throttle Burstable containers before they hit their hard limit. The Recent changes section covers the details.
Eviction: The kubelet's Triage
The kubelet watches node-level signals such as memory.available, nodefs.available, and pid.available. When one crosses an eviction threshold (for memory, a common hard threshold is memory.available<100Mi), the kubelet starts terminating pods to reclaim resources before the kernel is forced to act. Hard evictions use a zero-second grace period and ignore PodDisruptionBudgets.
Here is the part most articles get wrong: the kubelet does not rank pods by QoS class. For memory pressure, it ranks them by three criteria, in order:
- Whether the pod's memory usage exceeds its memory request.
- The pod's priority, from its
PriorityClass. - How far the pod's usage exceeds its request.
QoS falls out of those rules rather than driving them. A BestEffort pod has a request of zero, so any usage at all puts it over its request and into the first group. A Guaranteed pod's usage can never exceed its request, because its limit equals its request, so it never lands in that group. A Burstable pod is safe while it stays under its request and becomes a candidate the moment it goes above it.
That reframes the 3 a.m. incident. The payment service requested 512 MiB but routinely used 1.8 GiB. Under memory pressure, it was the pod with the largest overage on the node. The batch job requested exactly what it used, so it was never in the running. The kubelet did not think the batch job was more important; it thought the payment service was the one breaking its promise.
The OOM Killer: The Kernel's Triage
Eviction is proactive, and it runs on a polling loop. If memory spikes faster than the kubelet can react, the node hits real exhaustion first and the Linux OOM killer steps in. The kernel does not know about pods or priorities. It knows OOM scores, and each process's score is its memory footprint adjusted by oom_score_adj. This is where the kubelet writes QoS directly into the kernel.
Guaranteed containers get -997, which makes them nearly unkillable relative to anything else on the node. BestEffort containers get 1000, the maximum, so they are first in line. Burstable containers get a value computed from their memory request:
oom_score_adj = 1000 - (1000 × memory_request) / node_memory_capacity
clamped to the range [2, 999]
On a 16 GiB node, a Burstable container requesting 512 MiB gets 1000 - 31 = 969, nearly as exposed as BestEffort. One requesting 8 GiB gets 500. The larger your request relative to the node, the more protection you buy. In the 3 a.m. incident, the payment service's small request left it with a score in the high 900s, while the Guaranteed batch job sat at -997. The two triage systems reached the same answer by different routes.
Another subtlety: on cgroup v2, since v1.32, the kubelet sets memory.oom.group on containers. When the OOM killer picks any process in a container, it kills every process in that container. This avoids leaving a half-dead container running with its worker processes gone. Clusters that relied on the old single-process behavior can restore it with the kubelet setting singleProcessOOMKill.
Exclusive CPUs: The Guaranteed Perk
One capability belongs to Guaranteed pods alone. When a node runs the kubelet with the CPU Manager static policy, containers in Guaranteed pods that request a whole number of CPUs get exclusive cores. The kubelet pins them with cpuset.cpus, and every other container on the node is kept off those cores.
For latency-sensitive workloads such as packet processing, trading engines, and some databases, this removes both noisy neighbors and cross-core cache thrashing. It only applies to integer requests. A Guaranteed container with cpu: "1500m" stays in the shared pool. The Memory Manager offers a similar NUMA-aware guarantee for memory, again only for Guaranteed pods.
Commonly Confused Neighbors
PriorityClass is the one most often mixed up with QoS. Priority is set explicitly and is used mainly by the scheduler: when a high-priority pod cannot fit, the scheduler can preempt (evict) lower-priority pods to make room. QoS is derived and is used mainly by the node. The two meet in one place, the kubelet's eviction ranking, where priority is the tiebreaker after the "usage above request" check. A high-priority Burstable pod that exceeds its request can still be evicted before a low-priority pod that stays within its request.
ResourceQuota and LimitRange are namespace policies, not runtime behavior. ResourceQuota caps the total requests and limits in a namespace, and when it covers CPU or memory, it forces every pod to declare them, which rules out BestEffort. LimitRange injects default requests and limits, which can silently change a pod's class.
API-initiated eviction, the kind triggered by kubectl drain or the cluster autoscaler, has nothing to do with QoS at all. It respects PodDisruptionBudgets and graceful termination. Node-pressure eviction does neither.
Recent Changes
The core QoS rules have been stable for years, but the machinery around them has moved quickly. As of Kubernetes v1.37:
- In-place pod resize is stable since v1.35. You can change CPU and memory on a running pod through the
resizesubresource, but the QoS class cannot change. A resize that would turn a Guaranteed pod into a Burstable one is rejected, so pick the class you want to live with. - Pod-level resources are beta and enabled by default since v1.34. You can set requests and limits for the pod as a whole, and a pod-level request equal to a pod-level limit for both CPU and memory makes the pod Guaranteed.
- Memory QoS is beta and enabled by default since v1.37, but it does nothing until you configure the kubelet. Setting
memoryThrottlingFactorenablesmemory.highthrottling for Burstable and BestEffort containers. SettingmemoryReservationPolicy: TieredReservationsetsmemory.minto the request for Guaranteed pods andmemory.lowfor Burstable pods. The docs warn that for Guaranteed pods this makesmemory.minequal tomemory.max, so workloads with heavy page cache need headroom in their limit. - Whole-container OOM kills on cgroup v2, via
memory.oom.group, have been the default since v1.32.
Managed Kubernetes providers often lag the upstream release or change feature defaults, so check your provider's version and kubelet configuration before relying on any of these.
In Practice
Start by finding out what you actually have. The class is right there in the pod status:
# QoS class of every pod in a namespace
kubectl get pods -n payments \
-o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
# The OOM adjustment the kernel actually sees for a container's main process
kubectl exec -n payments payment-api-7f9c -- cat /proc/1/oom_score_adj
# CPU throttling counters (cgroup v2)
kubectl exec -n payments payment-api-7f9c -- cat /sys/fs/cgroup/cpu.stat
In cpu.stat, compare nr_throttled with nr_periods. If a meaningful fraction of periods are throttled while node CPU sits below capacity, your CPU limit is costing you latency for no benefit. The same data is exported by cAdvisor as container_cpu_cfs_throttled_periods_total, which is easier to graph over time.
Then make the decisions the mechanism implies:
- Set memory requests close to real usage. This single number drives eviction eligibility, your OOM score, and scheduling. An underestimated memory request is the most common reason an important Burstable pod dies first.
- Keep memory limits. They are the only thing stopping one leaking container from taking the whole node's memory down with it.
- Think carefully about CPU limits. Many teams drop them for latency-sensitive services and rely on requests for fair sharing, accepting Burstable class in exchange for no throttling. Others keep them for predictability or multi-tenant fairness. Either choice is defensible, as long as it is deliberate.
- Use Guaranteed on purpose. It is the right choice for workloads that must survive memory pressure or need exclusive cores, and a costly one for bursty workloads, since you reserve the peak permanently.
- Never run anything you care about as BestEffort. Reserve it for genuinely disposable work that should use spare capacity and nothing more.
Summary
QoS classes look like a label, but they are really a translation layer. The same two numbers you write in every manifest, requests and limits, are turned into a CFS weight, a CFS quota, a memory.max, a position in the cgroup tree, an eviction ranking, and an OOM score. Each of those takes effect under a different kind of pressure, which is why QoS behavior often seems inconsistent until you see all of it at once.
The insight that ties it together is this: a request is a promise, and pressure is when promises are audited. Pods that stay within what they asked for are protected, and pods that use more than they declared are the first to go, regardless of how important they are to your business. The 3 a.m. question is not "which service matters most?" but "which service told the node the truth about what it needs?" When you set requests and limits, you are writing the answer in advance.
Part of the Under the Hood series: what your systems actually do with your config. For the full picture of the components involved, see Architecture: Kubernetes.