Kubernetes requests vs limits and QoS (what actually happens)
Tabela de Conteúdo
If your pod is being OOMKilled or the app feels slow without a clear cause, the culprit is usually requests/limits.
The mental model is simple:
requests= reservation used by the scheduler to decide where to place the podlimits= ceiling the container cannot exceed (CPU can be throttled, memory can be OOMKilled)
QoS classes (what Kubernetes uses internally)
Guaranteed (requests == limits)
- Guaranteed resources
- No CPU throttling (within the limit)
- OOMKill if memory exceeds limit
- Criteria: every container must have requests == limits for both CPU and memory
Burstable (requests < limits)
- Can be preempted by Guaranteed pods
- CPU can be throttled under node pressure
- OOMKill if memory exceeds limit
- Criteria: at least one container with a request or limit set (but not all equal)
BestEffort (no requests/limits)
- First to be preempted if the node needs resources
- CPU can drop to near zero under pressure
- OOMKill if the node needs memory
- Criteria: no container has requests or limits set
Quick debug checklist
kubectl top pod -A
kubectl describe pod -n <ns> <pod>
kubectl describe node <node>
kubectl get events -n <ns> --sort-by=.lastTimestamp | tail -n 30
Look for:
Killing container with id ...(OOM)Throttling(CPU)Insufficient cpu/memory(preemption)
Common mistakes
- Setting only limits and ending up as BestEffort by accident
- Setting requests too high and wasting cluster resources
- Not setting memory limits and getting OOMKilled at the node level
- CPU limits too low, making the app feel slow
Copy/paste example (Burstable)
apiVersion: apps/v1
kind: Deployment
metadata:
name: api
spec:
replicas: 2
selector:
matchLabels:
app: api
template:
metadata:
labels:
app: api
spec:
containers:
- name: api
image: nginx:1.27
resources:
requests:
memory: "128Mi"
cpu: "100m"
limits:
memory: "256Mi"
cpu: "500m"
How to read this YAML
requests.memory: reserves 128 MiB (scheduler uses this for placement)limits.memory: ceiling of 256 MiB (OOMKill if exceeded)requests.cpu: reserves 0.1 vCPU (100 millicores)limits.cpu: ceiling of 0.5 vCPU (throttling if exceeded)
When to use each class
Guaranteed
- Critical apps that need guaranteed performance
- Services with predictable load
- When you can afford to pay for guaranteed resources
Burstable
- Most web applications
- Services with variable load
- When you want flexibility with some control
BestEffort
- Batch jobs that can be restarted
- Low-priority apps
- When resources aren’t critical
Useful commands
# Check the QoS class of your pods
kubectl get pod -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
# Check resource allocation per node
kubectl describe node <node> | grep -A 10 "Allocated resources"
# Check pods consuming more than requested
kubectl top pod -n <ns> --sort-by=cpu
kubectl top pod -n <ns> --sort-by=memory
# Check OOM events
kubectl get events -A --sort-by=.lastTimestamp | grep "Killing"
Best practices
- Always set requests if you set limits (otherwise you might get BestEffort unintentionally)
- For critical apps, consider Guaranteed (requests == limits)
- For batch or low-priority apps, Burstable is usually enough
- Monitor real usage and adjust as needed
If you want to understand how rollout and probes interact with resources, check out this post: Kubernetes probes: liveness vs readiness vs startup (practical).
Easy peasy! :)