Files
memory/patterns/pve-oom-frozen-guest.md
T

2.1 KiB

pattern_id, title, category, severity, status, first_observed, last_updated, related_systems, related_solution_docs, related_skills
pattern_id title category severity status first_observed last_updated related_systems related_solution_docs related_skills
PAT-010 PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature) infrastructure high active 2026-09 2026-09-30
proxmox-cluster
rke2-kubernetes
docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md
rke2-cluster-administration
systematic-debugging

PAT-010: PVE host OOM-kill freezes guest VM as zombie

Symptom

  • K8s node NotReady, kubelet heartbeat stops abruptly
  • PVE shows VM running, but: no ping, no SSH, qemu-guest-agent dead, tap-interface TX counters static
  • Signature: qm status <vmid> --verbose → kvm RSS collapses to a tiny fraction (<5%) of assigned RAM — guest kernel no longer touches its memory
  • Downstream alerts (DS misscheduled, workload CrashLoops) fire, but NOTHING alarms on the actual OOM

Root Cause

Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks the largest anon-RSS process = biggest kvm. After TWO consecutive kills of the same guest (HA auto-restarts in between), the revived QEMU comes up but the guest kernel stays inert → zombie VM.

Mitigation

  1. Relieve host pressure FIRST (move movable tenants away) — otherwise the unfreeze re-boots the guest into the same thrash.
  2. Hard cycle the VM: qm stop + qm start (respect HA guards).
  3. Re-query ha-manager status afterwards — HA may relocate the VM.
  4. Uncordon the K8s node; orphan DS pods self-reconcile.

Prevention

  • Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/ IaC; swap is a buffer, not capacity.
  • PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently invisible to alerting (noticed only via downstream K8s symptoms).

Evidence

  • 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration + hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.