PVE shows VM running, but: no ping, no SSH, qemu-guest-agent dead,
tap-interface TX counters static
Signature:qm status <vmid> --verbose → kvm RSS collapses to a
tiny fraction (<5%) of assigned RAM — guest kernel no longer touches
its memory
Downstream alerts (DS misscheduled, workload CrashLoops) fire, but
NOTHING alarms on the actual OOM
Root Cause
Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks
the largest anon-RSS process = biggest kvm. After TWO consecutive kills of
the same guest (HA auto-restarts in between), the revived QEMU comes up
but the guest kernel stays inert → zombie VM.
Mitigation
Relieve host pressure FIRST (move movable tenants away) — otherwise
the unfreeze re-boots the guest into the same thrash.
Hard cycle the VM: qm stop + qm start (respect HA guards).
Re-query ha-manager status afterwards — HA may relocate the VM.
Uncordon the K8s node; orphan DS pods self-reconcile.
Prevention
Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/
IaC; swap is a buffer, not capacity.
PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently
invisible to alerting (noticed only via downstream K8s symptoms).
Evidence
2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed
rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second
revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration +
hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.