ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping

This commit is contained in:
Dominik Schön
2026-09-30 18:22:25 +00:00
parent 531d6bfbdd
commit cc2f237ca3
8 changed files with 183 additions and 8 deletions
+14 -7
View File
@@ -50,17 +50,24 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
- Benötigte modprobe.d Config:
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
- Worker-05 (VM 102): GPU-Passthrough funktionierte historisch auf ms-a2-2 (identische Hardware). 2026-09-29 stand VM auf proxmox6 (OHNE GPU) wegen Anti-Collocation; seit 2026-09-30 hat HA die VM nach ms-a2-2 zurückplatziert (nach OOM-Freezer-Incident, siehe PAT-010). GPU-Status nach Return prüfen/renderD128 verifizieren. Hinweis: Die frühere node-affinity-Regel `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr — nur noch resource-affinity (anti-colloc vs 128/139).
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
## Bekannte Probleme
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
- **proxmox6 RAM-Oversubscription (STRUKTURELL):** ~28 GB Gast-Allokation
(VM301 8G + CT151 Frigate 8G [+ VM102-Rückkehr möglich 12G]) auf 16 GB
Host. Führte 2026-09-29/30 zu Doppel-OOM-Kill von VM102 (Zombie-VM,
PAT-010). Akut entschärft durch Laya-CT152-Migration nach proxmox3.
TODO: Placement-Ceiling (~70% RAM) + PSI/OOM-Alerts pro Node (CT141).
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
## HA Rules (PVE 9.2 Rules System)
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
@@ -119,7 +126,7 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
| Hermes-Agent-01 | 230 | n5pro |
| Galera db1 | 300 | n5pro |
| Galera db2 | 301 | proxmox6 |
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
| Galera db3 | 302 | ms-a2-2 |
| MaxScale-01 | 310 | proxmox7 |
| MaxScale-02 | 311 | ms-a2-1 |