Kubernetes requests vs limits and QoS (what actually happens)

Tabela de Conteúdo

If your pod is being OOMKilled or the app feels slow without a clear cause, the culprit is usually requests/limits.

The mental model is simple:

  • requests = reservation used by the scheduler to decide where to place the pod
  • limits = ceiling the container cannot exceed (CPU can be throttled, memory can be OOMKilled)

QoS classes (what Kubernetes uses internally)

Guaranteed (requests == limits)

  • Guaranteed resources
  • No CPU throttling (within the limit)
  • OOMKill if memory exceeds limit
  • Criteria: every container must have requests == limits for both CPU and memory

Burstable (requests < limits)

  • Can be preempted by Guaranteed pods
  • CPU can be throttled under node pressure
  • OOMKill if memory exceeds limit
  • Criteria: at least one container with a request or limit set (but not all equal)

BestEffort (no requests/limits)

  • First to be preempted if the node needs resources
  • CPU can drop to near zero under pressure
  • OOMKill if the node needs memory
  • Criteria: no container has requests or limits set

Quick debug checklist

kubectl top pod -A
kubectl describe pod -n <ns> <pod>
kubectl describe node <node>
kubectl get events -n <ns> --sort-by=.lastTimestamp | tail -n 30

Look for:

  • Killing container with id ... (OOM)
  • Throttling (CPU)
  • Insufficient cpu/memory (preemption)

Common mistakes

  • Setting only limits and ending up as BestEffort by accident
  • Setting requests too high and wasting cluster resources
  • Not setting memory limits and getting OOMKilled at the node level
  • CPU limits too low, making the app feel slow

Copy/paste example (Burstable)

apiVersion: apps/v1
kind: Deployment
metadata:
  name: api
spec:
  replicas: 2
  selector:
    matchLabels:
      app: api
  template:
    metadata:
      labels:
        app: api
    spec:
      containers:
        - name: api
          image: nginx:1.27
          resources:
            requests:
              memory: "128Mi"
              cpu: "100m"
            limits:
              memory: "256Mi"
              cpu: "500m"

How to read this YAML

  • requests.memory: reserves 128 MiB (scheduler uses this for placement)
  • limits.memory: ceiling of 256 MiB (OOMKill if exceeded)
  • requests.cpu: reserves 0.1 vCPU (100 millicores)
  • limits.cpu: ceiling of 0.5 vCPU (throttling if exceeded)

When to use each class

Guaranteed

  • Critical apps that need guaranteed performance
  • Services with predictable load
  • When you can afford to pay for guaranteed resources

Burstable

  • Most web applications
  • Services with variable load
  • When you want flexibility with some control

BestEffort

  • Batch jobs that can be restarted
  • Low-priority apps
  • When resources aren’t critical

Useful commands

# Check the QoS class of your pods
kubectl get pod -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass

# Check resource allocation per node
kubectl describe node <node> | grep -A 10 "Allocated resources"

# Check pods consuming more than requested
kubectl top pod -n <ns> --sort-by=cpu
kubectl top pod -n <ns> --sort-by=memory

# Check OOM events
kubectl get events -A --sort-by=.lastTimestamp | grep "Killing"

Best practices

  • Always set requests if you set limits (otherwise you might get BestEffort unintentionally)
  • For critical apps, consider Guaranteed (requests == limits)
  • For batch or low-priority apps, Burstable is usually enough
  • Monitor real usage and adjust as needed

If you want to understand how rollout and probes interact with resources, check out this post: Kubernetes probes: liveness vs readiness vs startup (practical).

Easy peasy! :)


References