ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping

This commit is contained in:
Dominik Schön
2026-09-30 18:22:25 +00:00
parent 531d6bfbdd
commit cc2f237ca3
8 changed files with 183 additions and 8 deletions
+13 -1
View File
@@ -23,7 +23,7 @@ modified: "2026-09-26"
|-----|-------|------|------|----------|-------|
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
| 3 | ssd | 233 GB | proxmox4 | 1.0 | |
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
@@ -73,6 +73,18 @@ Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not al
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
Consider removing from CRUSH or replacing with larger drives.
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
thermal throttle T1 3×.**
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
### worker-04 (VM 139) NotReady in K8s
Node offline — not a Ceph issue but affects Ceph CSI attachments.
+14 -7
View File
@@ -50,17 +50,24 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
- Benötigte modprobe.d Config:
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
- Worker-05 (VM 102): GPU-Passthrough funktionierte historisch auf ms-a2-2 (identische Hardware). 2026-09-29 stand VM auf proxmox6 (OHNE GPU) wegen Anti-Collocation; seit 2026-09-30 hat HA die VM nach ms-a2-2 zurückplatziert (nach OOM-Freezer-Incident, siehe PAT-010). GPU-Status nach Return prüfen/renderD128 verifizieren. Hinweis: Die frühere node-affinity-Regel `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr — nur noch resource-affinity (anti-colloc vs 128/139).
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
## Bekannte Probleme
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
- **proxmox6 RAM-Oversubscription (STRUKTURELL):** ~28 GB Gast-Allokation
(VM301 8G + CT151 Frigate 8G [+ VM102-Rückkehr möglich 12G]) auf 16 GB
Host. Führte 2026-09-29/30 zu Doppel-OOM-Kill von VM102 (Zombie-VM,
PAT-010). Akut entschärft durch Laya-CT152-Migration nach proxmox3.
TODO: Placement-Ceiling (~70% RAM) + PSI/OOM-Alerts pro Node (CT141).
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
## HA Rules (PVE 9.2 Rules System)
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
@@ -119,7 +126,7 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
| Hermes-Agent-01 | 230 | n5pro |
| Galera db1 | 300 | n5pro |
| Galera db2 | 301 | proxmox6 |
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
| Galera db3 | 302 | ms-a2-2 |
| MaxScale-01 | 310 | proxmox7 |
| MaxScale-02 | 311 | ms-a2-1 |