From 531d6bfbddda4ddb8faacca61b852971b3f27ae2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Dominik=20Sch=C3=B6n?= Date: Wed, 30 Sep 2026 14:15:18 +0000 Subject: [PATCH] docs(wiki): PAT-010 pve-oom-frozen-guest + proxmox6 drift fixes (worker-05 placement, oversubscription known-issue) --- index.md | 1 + log.md | 9 ++++++ patterns/pve-oom-frozen-guest.md | 49 ++++++++++++++++++++++++++++++++ systems/proxmox-cluster.md | 12 +++++--- systems/rke2-kubernetes.md | 2 +- 5 files changed, 68 insertions(+), 5 deletions(-) create mode 100644 patterns/pve-oom-frozen-guest.md diff --git a/index.md b/index.md index 79b44aa..4ddc046 100644 --- a/index.md +++ b/index.md @@ -61,6 +61,7 @@ - [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007) - [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008) - [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009) +- [[patterns/pve-oom-frozen-guest]] — Host-OOM friert Gast als Zombie ein, RSS-Kollaps-Signatur (PAT-010) - [[patterns/skill-impact]] — Skill Modification Audit Trail ## Memory Layer Architektur diff --git a/log.md b/log.md index 9fdc7ff..a31847c 100644 --- a/log.md +++ b/log.md @@ -1,5 +1,14 @@ # Memory Log +## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010) +- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h +- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB) +- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX +- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared +- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr +- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return +- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010) + ## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live - VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil) - hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key) diff --git a/patterns/pve-oom-frozen-guest.md b/patterns/pve-oom-frozen-guest.md new file mode 100644 index 0000000..405a711 --- /dev/null +++ b/patterns/pve-oom-frozen-guest.md @@ -0,0 +1,49 @@ +--- +pattern_id: PAT-010 +title: "PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature)" +category: infrastructure +severity: high +status: active +first_observed: 2026-09 +last_updated: 2026-09-30 +related_systems: [proxmox-cluster, rke2-kubernetes] +related_solution_docs: [docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md] +related_skills: [rke2-cluster-administration, systematic-debugging] +--- + +# PAT-010: PVE host OOM-kill freezes guest VM as zombie + +## Symptom +- K8s node NotReady, kubelet heartbeat stops abruptly +- PVE shows VM `running`, but: no ping, no SSH, qemu-guest-agent dead, + tap-interface TX counters static +- **Signature:** `qm status --verbose` → kvm RSS collapses to a + tiny fraction (<5%) of assigned RAM — guest kernel no longer touches + its memory +- Downstream alerts (DS misscheduled, workload CrashLoops) fire, but + NOTHING alarms on the actual OOM + +## Root Cause +Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks +the largest anon-RSS process = biggest kvm. After TWO consecutive kills of +the same guest (HA auto-restarts in between), the revived QEMU comes up +but the guest kernel stays inert → zombie VM. + +## Mitigation +1. Relieve host pressure FIRST (move movable tenants away) — otherwise + the unfreeze re-boots the guest into the same thrash. +2. Hard cycle the VM: `qm stop` + `qm start` (respect HA guards). +3. Re-query `ha-manager status` afterwards — HA may relocate the VM. +4. Uncordon the K8s node; orphan DS pods self-reconcile. + +## Prevention +- Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/ + IaC; swap is a buffer, not capacity. +- PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently + invisible to alerting (noticed only via downstream K8s symptoms). + +## Evidence +- 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed + rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second + revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration + + hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc. diff --git a/systems/proxmox-cluster.md b/systems/proxmox-cluster.md index f9425fb..2087818 100644 --- a/systems/proxmox-cluster.md +++ b/systems/proxmox-cluster.md @@ -3,7 +3,7 @@ title: Proxmox VE Cluster category: systems tags: [proxmox, virtualization, lxc, qemu, pve] created: "2026-04-28" -modified: "2026-09-29" +modified: "2026-09-30" --- # Proxmox VE Cluster @@ -50,13 +50,17 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs - Benötigte modprobe.d Config: - `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper` - `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci` -- Worker-05 (VM 102) lief auf ms-a2-2 mit funktionierendem GPU-Passthrough. Seit 2026-09-29 auf proxmox6 (Anti-Collocation mit Worker-01). Prüfen ob GPU-Passthrough auf proxmox6 funktioniert! +- Worker-05 (VM 102): GPU-Passthrough funktionierte historisch auf ms-a2-2 (identische Hardware). 2026-09-29 stand VM auf proxmox6 (OHNE GPU) wegen Anti-Collocation; seit 2026-09-30 hat HA die VM nach ms-a2-2 zurückplatziert (nach OOM-Freezer-Incident, siehe PAT-010). GPU-Status nach Return prüfen/renderD128 verifizieren. Hinweis: Die frühere node-affinity-Regel `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr — nur noch resource-affinity (anti-colloc vs 128/139). ## Bekannte Probleme - CT110 kaputte libc — Reparatur ausstehend (still stopped) - osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen - osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation -- worker-04 (VM 139) NotReady im K8s Cluster — Node offline +- **proxmox6 RAM-Oversubscription (STRUKTURELL):** ~28 GB Gast-Allokation + (VM301 8G + CT151 Frigate 8G [+ VM102-Rückkehr möglich 12G]) auf 16 GB + Host. Führte 2026-09-29/30 zu Doppel-OOM-Kill von VM102 (Zombie-VM, + PAT-010). Akut entschärft durch Laya-CT152-Migration nach proxmox3. + TODO: Placement-Ceiling (~70% RAM) + PSI/OOM-Alerts pro Node (CT141). ## HA Rules (PVE 9.2 Rules System) Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity. @@ -112,7 +116,7 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt. | RKE2 CP-03 | 126 | proxmox4 | | RKE2 Worker-01 | 128 | proxmox5 | | RKE2 Worker-04 | 139 | n5pro | -| RKE2 Worker-05 | 102 | proxmox6 | +| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) | | Hermes-Agent-01 | 230 | n5pro | | Galera db1 | 300 | n5pro | | Galera db2 | 301 | proxmox6 | diff --git a/systems/rke2-kubernetes.md b/systems/rke2-kubernetes.md index 50dd0ee..b5e29bc 100644 --- a/systems/rke2-kubernetes.md +++ b/systems/rke2-kubernetes.md @@ -22,7 +22,7 @@ modified: "2026-09-17" | cp-03 | 10.0.30.53 | Control Plane | | worker-01 | 10.0.30.63 | Worker | | worker-04 | 10.0.30.64 | Worker | -| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) | +| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2; 30.09. OOM-Zombie-Freeze behoben, siehe PAT-010) | ## Storage - **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec