ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping

This commit is contained in:
Dominik Schön
2026-09-30 18:22:25 +00:00
parent 531d6bfbdd
commit cc2f237ca3
8 changed files with 183 additions and 8 deletions
+13 -1
View File
@@ -23,7 +23,7 @@ modified: "2026-09-26"
|-----|-------|------|------|----------|-------|
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
| 3 | ssd | 233 GB | proxmox4 | 1.0 | |
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
@@ -73,6 +73,18 @@ Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not al
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
Consider removing from CRUSH or replacing with larger drives.
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
thermal throttle T1 3×.**
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
### worker-04 (VM 139) NotReady in K8s
Node offline — not a Ceph issue but affects Ceph CSI attachments.