ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping
This commit is contained in:
+13
-1
@@ -23,7 +23,7 @@ modified: "2026-09-26"
|
||||
|-----|-------|------|------|----------|-------|
|
||||
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
|
||||
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
|
||||
| 3 | ssd | 233 GB | proxmox4 | 1.0 | |
|
||||
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
|
||||
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
|
||||
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
|
||||
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
|
||||
@@ -73,6 +73,18 @@ Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not al
|
||||
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
|
||||
Consider removing from CRUSH or replacing with larger drives.
|
||||
|
||||
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
|
||||
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
|
||||
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
|
||||
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
|
||||
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
|
||||
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
|
||||
thermal throttle T1 3×.**
|
||||
|
||||
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
|
||||
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
|
||||
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
|
||||
|
||||
### worker-04 (VM 139) NotReady in K8s
|
||||
Node offline — not a Ceph issue but affects Ceph CSI attachments.
|
||||
|
||||
|
||||
@@ -50,17 +50,24 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
||||
- Benötigte modprobe.d Config:
|
||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||
- Worker-05 (VM 102): GPU-Passthrough funktionierte historisch auf ms-a2-2 (identische Hardware). 2026-09-29 stand VM auf proxmox6 (OHNE GPU) wegen Anti-Collocation; seit 2026-09-30 hat HA die VM nach ms-a2-2 zurückplatziert (nach OOM-Freezer-Incident, siehe PAT-010). GPU-Status nach Return prüfen/renderD128 verifizieren. Hinweis: Die frühere node-affinity-Regel `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr — nur noch resource-affinity (anti-colloc vs 128/139).
|
||||
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
|
||||
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
|
||||
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
|
||||
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
|
||||
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
|
||||
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
|
||||
|
||||
## Bekannte Probleme
|
||||
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
||||
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
||||
- **proxmox6 RAM-Oversubscription (STRUKTURELL):** ~28 GB Gast-Allokation
|
||||
(VM301 8G + CT151 Frigate 8G [+ VM102-Rückkehr möglich 12G]) auf 16 GB
|
||||
Host. Führte 2026-09-29/30 zu Doppel-OOM-Kill von VM102 (Zombie-VM,
|
||||
PAT-010). Akut entschärft durch Laya-CT152-Migration nach proxmox3.
|
||||
TODO: Placement-Ceiling (~70% RAM) + PSI/OOM-Alerts pro Node (CT141).
|
||||
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
|
||||
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
|
||||
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
|
||||
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
|
||||
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
|
||||
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
|
||||
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
|
||||
|
||||
## HA Rules (PVE 9.2 Rules System)
|
||||
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||
@@ -119,7 +126,7 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
||||
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
||||
| Hermes-Agent-01 | 230 | n5pro |
|
||||
| Galera db1 | 300 | n5pro |
|
||||
| Galera db2 | 301 | proxmox6 |
|
||||
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
|
||||
| Galera db3 | 302 | ms-a2-2 |
|
||||
| MaxScale-01 | 310 | proxmox7 |
|
||||
| MaxScale-02 | 311 | ms-a2-1 |
|
||||
|
||||
Reference in New Issue
Block a user