ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping
This commit is contained in:
@@ -50,17 +50,24 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
||||
- Benötigte modprobe.d Config:
|
||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||
- Worker-05 (VM 102): GPU-Passthrough funktionierte historisch auf ms-a2-2 (identische Hardware). 2026-09-29 stand VM auf proxmox6 (OHNE GPU) wegen Anti-Collocation; seit 2026-09-30 hat HA die VM nach ms-a2-2 zurückplatziert (nach OOM-Freezer-Incident, siehe PAT-010). GPU-Status nach Return prüfen/renderD128 verifizieren. Hinweis: Die frühere node-affinity-Regel `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr — nur noch resource-affinity (anti-colloc vs 128/139).
|
||||
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
|
||||
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
|
||||
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
|
||||
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
|
||||
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
|
||||
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
|
||||
|
||||
## Bekannte Probleme
|
||||
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
||||
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
||||
- **proxmox6 RAM-Oversubscription (STRUKTURELL):** ~28 GB Gast-Allokation
|
||||
(VM301 8G + CT151 Frigate 8G [+ VM102-Rückkehr möglich 12G]) auf 16 GB
|
||||
Host. Führte 2026-09-29/30 zu Doppel-OOM-Kill von VM102 (Zombie-VM,
|
||||
PAT-010). Akut entschärft durch Laya-CT152-Migration nach proxmox3.
|
||||
TODO: Placement-Ceiling (~70% RAM) + PSI/OOM-Alerts pro Node (CT141).
|
||||
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
|
||||
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
|
||||
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
|
||||
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
|
||||
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
|
||||
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
|
||||
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
|
||||
|
||||
## HA Rules (PVE 9.2 Rules System)
|
||||
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||
@@ -119,7 +126,7 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
||||
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
||||
| Hermes-Agent-01 | 230 | n5pro |
|
||||
| Galera db1 | 300 | n5pro |
|
||||
| Galera db2 | 301 | proxmox6 |
|
||||
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
|
||||
| Galera db3 | 302 | ms-a2-2 |
|
||||
| MaxScale-01 | 310 | proxmox7 |
|
||||
| MaxScale-02 | 311 | ms-a2-1 |
|
||||
|
||||
Reference in New Issue
Block a user