docs(wiki): PAT-010 pve-oom-frozen-guest + proxmox6 drift fixes (worker-05 placement, oversubscription known-issue)
This commit is contained in:
@@ -61,6 +61,7 @@
|
|||||||
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
|
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
|
||||||
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
|
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
|
||||||
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
|
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
|
||||||
|
- [[patterns/pve-oom-frozen-guest]] — Host-OOM friert Gast als Zombie ein, RSS-Kollaps-Signatur (PAT-010)
|
||||||
- [[patterns/skill-impact]] — Skill Modification Audit Trail
|
- [[patterns/skill-impact]] — Skill Modification Audit Trail
|
||||||
|
|
||||||
## Memory Layer Architektur
|
## Memory Layer Architektur
|
||||||
|
|||||||
@@ -1,5 +1,14 @@
|
|||||||
# Memory Log
|
# Memory Log
|
||||||
|
|
||||||
|
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
|
||||||
|
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
|
||||||
|
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
|
||||||
|
- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX
|
||||||
|
- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared
|
||||||
|
- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr
|
||||||
|
- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return
|
||||||
|
- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010)
|
||||||
|
|
||||||
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
|
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
|
||||||
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
|
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
|
||||||
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
|
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
|
||||||
|
|||||||
@@ -0,0 +1,49 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-010
|
||||||
|
title: "PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature)"
|
||||||
|
category: infrastructure
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-09
|
||||||
|
last_updated: 2026-09-30
|
||||||
|
related_systems: [proxmox-cluster, rke2-kubernetes]
|
||||||
|
related_solution_docs: [docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md]
|
||||||
|
related_skills: [rke2-cluster-administration, systematic-debugging]
|
||||||
|
---
|
||||||
|
|
||||||
|
# PAT-010: PVE host OOM-kill freezes guest VM as zombie
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
- K8s node NotReady, kubelet heartbeat stops abruptly
|
||||||
|
- PVE shows VM `running`, but: no ping, no SSH, qemu-guest-agent dead,
|
||||||
|
tap-interface TX counters static
|
||||||
|
- **Signature:** `qm status <vmid> --verbose` → kvm RSS collapses to a
|
||||||
|
tiny fraction (<5%) of assigned RAM — guest kernel no longer touches
|
||||||
|
its memory
|
||||||
|
- Downstream alerts (DS misscheduled, workload CrashLoops) fire, but
|
||||||
|
NOTHING alarms on the actual OOM
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks
|
||||||
|
the largest anon-RSS process = biggest kvm. After TWO consecutive kills of
|
||||||
|
the same guest (HA auto-restarts in between), the revived QEMU comes up
|
||||||
|
but the guest kernel stays inert → zombie VM.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
1. Relieve host pressure FIRST (move movable tenants away) — otherwise
|
||||||
|
the unfreeze re-boots the guest into the same thrash.
|
||||||
|
2. Hard cycle the VM: `qm stop` + `qm start` (respect HA guards).
|
||||||
|
3. Re-query `ha-manager status` afterwards — HA may relocate the VM.
|
||||||
|
4. Uncordon the K8s node; orphan DS pods self-reconcile.
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
- Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/
|
||||||
|
IaC; swap is a buffer, not capacity.
|
||||||
|
- PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently
|
||||||
|
invisible to alerting (noticed only via downstream K8s symptoms).
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
- 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed
|
||||||
|
rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second
|
||||||
|
revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration +
|
||||||
|
hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.
|
||||||
@@ -3,7 +3,7 @@ title: Proxmox VE Cluster
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [proxmox, virtualization, lxc, qemu, pve]
|
tags: [proxmox, virtualization, lxc, qemu, pve]
|
||||||
created: "2026-04-28"
|
created: "2026-04-28"
|
||||||
modified: "2026-09-29"
|
modified: "2026-09-30"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Proxmox VE Cluster
|
# Proxmox VE Cluster
|
||||||
@@ -50,13 +50,17 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
|||||||
- Benötigte modprobe.d Config:
|
- Benötigte modprobe.d Config:
|
||||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||||
- Worker-05 (VM 102) lief auf ms-a2-2 mit funktionierendem GPU-Passthrough. Seit 2026-09-29 auf proxmox6 (Anti-Collocation mit Worker-01). Prüfen ob GPU-Passthrough auf proxmox6 funktioniert!
|
- Worker-05 (VM 102): GPU-Passthrough funktionierte historisch auf ms-a2-2 (identische Hardware). 2026-09-29 stand VM auf proxmox6 (OHNE GPU) wegen Anti-Collocation; seit 2026-09-30 hat HA die VM nach ms-a2-2 zurückplatziert (nach OOM-Freezer-Incident, siehe PAT-010). GPU-Status nach Return prüfen/renderD128 verifizieren. Hinweis: Die frühere node-affinity-Regel `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr — nur noch resource-affinity (anti-colloc vs 128/139).
|
||||||
|
|
||||||
## Bekannte Probleme
|
## Bekannte Probleme
|
||||||
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||||
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
||||||
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
||||||
- worker-04 (VM 139) NotReady im K8s Cluster — Node offline
|
- **proxmox6 RAM-Oversubscription (STRUKTURELL):** ~28 GB Gast-Allokation
|
||||||
|
(VM301 8G + CT151 Frigate 8G [+ VM102-Rückkehr möglich 12G]) auf 16 GB
|
||||||
|
Host. Führte 2026-09-29/30 zu Doppel-OOM-Kill von VM102 (Zombie-VM,
|
||||||
|
PAT-010). Akut entschärft durch Laya-CT152-Migration nach proxmox3.
|
||||||
|
TODO: Placement-Ceiling (~70% RAM) + PSI/OOM-Alerts pro Node (CT141).
|
||||||
|
|
||||||
## HA Rules (PVE 9.2 Rules System)
|
## HA Rules (PVE 9.2 Rules System)
|
||||||
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||||
@@ -112,7 +116,7 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
|||||||
| RKE2 CP-03 | 126 | proxmox4 |
|
| RKE2 CP-03 | 126 | proxmox4 |
|
||||||
| RKE2 Worker-01 | 128 | proxmox5 |
|
| RKE2 Worker-01 | 128 | proxmox5 |
|
||||||
| RKE2 Worker-04 | 139 | n5pro |
|
| RKE2 Worker-04 | 139 | n5pro |
|
||||||
| RKE2 Worker-05 | 102 | proxmox6 |
|
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
||||||
| Hermes-Agent-01 | 230 | n5pro |
|
| Hermes-Agent-01 | 230 | n5pro |
|
||||||
| Galera db1 | 300 | n5pro |
|
| Galera db1 | 300 | n5pro |
|
||||||
| Galera db2 | 301 | proxmox6 |
|
| Galera db2 | 301 | proxmox6 |
|
||||||
|
|||||||
@@ -22,7 +22,7 @@ modified: "2026-09-17"
|
|||||||
| cp-03 | 10.0.30.53 | Control Plane |
|
| cp-03 | 10.0.30.53 | Control Plane |
|
||||||
| worker-01 | 10.0.30.63 | Worker |
|
| worker-01 | 10.0.30.63 | Worker |
|
||||||
| worker-04 | 10.0.30.64 | Worker |
|
| worker-04 | 10.0.30.64 | Worker |
|
||||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2; 30.09. OOM-Zombie-Freeze behoben, siehe PAT-010) |
|
||||||
|
|
||||||
## Storage
|
## Storage
|
||||||
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
||||||
|
|||||||
Reference in New Issue
Block a user