ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping
This commit is contained in:
@@ -1,5 +1,40 @@
|
||||
# Memory Log
|
||||
|
||||
## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014)
|
||||
- osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin.
|
||||
- osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend.
|
||||
- ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben.
|
||||
- Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014.
|
||||
|
||||
## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013)
|
||||
- Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün.
|
||||
- Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok.
|
||||
- **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate.
|
||||
- Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit <SHA>.
|
||||
|
||||
## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012)
|
||||
- Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok.
|
||||
- Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom).
|
||||
- Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172.
|
||||
|
||||
## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011)
|
||||
- Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON).
|
||||
- Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200.
|
||||
- Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body.
|
||||
- Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales).
|
||||
|
||||
## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen
|
||||
- proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%.
|
||||
- Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets.
|
||||
- VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP.
|
||||
- Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt.
|
||||
|
||||
## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet
|
||||
- GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md.
|
||||
- Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc.
|
||||
- proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen.
|
||||
- Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert.
|
||||
|
||||
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
|
||||
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
|
||||
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
|
||||
|
||||
Reference in New Issue
Block a user