From 02f6a8e2b0371388b871e2d95a80e383ec87e088 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Dominik=20Sch=C3=B6n?= Date: Sat, 26 Sep 2026 22:00:22 +0000 Subject: [PATCH] Auto-sync: 2026-09-26 --- entities/infrastructure.md | 11 +++-- index.md | 4 +- reference/ip-map.md | 31 +++++++------ systems/ceph-cluster.md | 91 ++++++++++++++++---------------------- systems/proxmox-cluster.md | 14 +++--- 5 files changed, 68 insertions(+), 83 deletions(-) diff --git a/entities/infrastructure.md b/entities/infrastructure.md index 93e7afb..723ff5b 100644 --- a/entities/infrastructure.md +++ b/entities/infrastructure.md @@ -3,24 +3,27 @@ title: Infrastruktur-Übersicht category: entities tags: [homelab, hardware, overview] created: "2026-04-28" -modified: "2026-07-24" +modified: "2026-09-26" --- # Infrastruktur-Übersicht ## Physikalische Hardware -### Proxmox Cluster (PVE 9.2.3) -- **9 Nodes**, Quorum OK +### Proxmox Cluster (PVE 9.2.20) +- **8 Nodes**, Quorum OK (proxmox1+2 dauerhaft entfernt) - Hypervisoren in 10.0.20.x - Siehe [[systems/proxmox-cluster]] ### Ceph Cluster -- **14 OSDs** (HDD + NVMe/SSD混合) +- **17 OSDs** (10 HDD + 7 SSD), all up/in, v20.2.4 +- HEALTH_WARN (insecure key types — cosmetic) +- 4 MONs (px5 leader, px4, a2-1, n5pro), MGR auf n5pro - Siehe [[systems/ceph-cluster]] ### RKE2 Kubernetes Cluster - **6 Nodes** (3 CP + 3 Worker), v1.35.6-rke2r1 +- worker-04 currently NotReady (VM 139 offline) - Alle Nodes schedulable (keine CP Taints) - Siehe [[systems/rke2-kubernetes]] diff --git a/index.md b/index.md index dcde658..2eea509 100644 --- a/index.md +++ b/index.md @@ -22,8 +22,8 @@ - [[entities/health-fitness]] — Gesundheits-Ziele, Ernährung ## Systems -- [[systems/proxmox-cluster]] — PVE 9.2.3, 9 Nodes, Fluent Bit -- [[systems/ceph-cluster]] — 14 OSDs, EC Pools, bekannte Probleme +- [[systems/proxmox-cluster]] — PVE 9.2.20, 8 Nodes, Fluent Bit +- [[systems/ceph-cluster]] — 17 OSDs, v20.2.4, 4 MONs, bekannte Probleme - [[systems/rke2-kubernetes]] — 6 Nodes, ArgoCD, CNPG, Workloads - [[systems/galera-maxscale]] — 3 Galera + 2 MaxScale, VIP .70 - [[systems/loki-fluentbit]] — Logging Stack, 40+ Targets diff --git a/reference/ip-map.md b/reference/ip-map.md index ee6ba36..8164832 100644 --- a/reference/ip-map.md +++ b/reference/ip-map.md @@ -3,37 +3,36 @@ title: IP-Map (Quick Reference) category: reference tags: [ip, network, reference, quick-lookup] created: "2026-07-24" -modified: "2026-07-24" +modified: "2026-09-26" --- # IP-Map -## Proxmox Hosts (10.0.20.x) — 9 Nodes +## Proxmox Hosts (10.0.20.x) — 8 Nodes | IP | Hostname | Node ID | Notes | |----|----------|---------|-------| -| 10.0.20.20 | proxmox2 | 3 | | | 10.0.20.30 | proxmox3 | 4 | | -| 10.0.20.40 | proxmox4 | 2 | | -| 10.0.20.50 | proxmox5 | 5 | MON, MGR | +| 10.0.20.40 | proxmox4 | 2 | MON | +| 10.0.20.50 | proxmox5 | 5 | MON (leader) | | 10.0.20.60 | proxmox6 | 6 | | | 10.0.20.70 | proxmox7 | 7 | MON, Traefik CT99999 | -| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002 | -| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0 | +| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002, MON, MGR (active) | +| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0, MON | | 10.0.20.93 | ms-a2-2 | 1 | OSD 14 (SSD 1.8TB), GPU 1002:13c0 | > ⚠️ n5pro (10.0.20.91, nodeid 8) ≠ proxmox7 (10.0.20.70, nodeid 7) — separate Nodes! -> proxmox1 (10.0.20.10, nodeid 1) wurde dauerhaft entfernt (2026-07-24). +> proxmox1 (10.0.20.10) und proxmox2 (10.0.20.20) wurden dauerhaft entfernt. > Immer `pvecm nodes` für kanonische Liste prüfen. ## Kubernetes Nodes (10.0.30.5x-6x) -| IP | Node | Role | -|----|------|------| -| 10.0.30.51 | cp-01 | Control Plane | -| 10.0.30.52 | cp-02 | Control Plane | -| 10.0.30.53 | cp-03 | Control Plane | -| 10.0.30.63 | worker-01 | Worker | -| 10.0.30.64 | worker-04 | Worker (GPU ✅) | -| 10.0.30.65 | worker-05 | Worker (GPU defekt) | +| IP | Node | Role | Status | +|----|------|------|--------| +| 10.0.30.51 | cp-01 | Control Plane | Ready | +| 10.0.30.52 | cp-02 | Control Plane | Ready | +| 10.0.30.53 | cp-03 | Control Plane | Ready | +| 10.0.30.61 | worker-01 | Worker | Ready | +| 10.0.30.64 | worker-04 | Worker (GPU) | **NotReady** | +| 10.0.30.65 | worker-05 | Worker (GPU) | Ready | ## Database Layer (10.0.30.7x-8x) | IP | Host | Service | diff --git a/systems/ceph-cluster.md b/systems/ceph-cluster.md index 29db407..a2736fe 100644 --- a/systems/ceph-cluster.md +++ b/systems/ceph-cluster.md @@ -3,37 +3,44 @@ title: Ceph Cluster category: systems tags: [ceph, storage, rbd, ec-pool, osd] created: "2026-07-24" -modified: "2026-07-25" +modified: "2026-09-26" --- # Ceph Cluster ## Overview - **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9 -- **Health**: HEALTH_OK (recovery complete after OSD 0+2 drain, 0.3% misplaced settling) -- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH 2026-07-25), 3 MONs (proxmox5/7/4), MGR on proxmox5 -- **OSDs**: 13 (8 SSD, 5 HDD), all up/in — OSDs 0+2 destroyed+purged 2026-07-25 -- **Capacity**: ~22 TiB total, 6.0 TiB used +- **Health**: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed) +- **Version**: 20.2.4 (tentacle) — all 17 OSDs +- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1) +- **OSDs**: 17 (10 HDD, 7 SSD), all up/in +- **Capacity**: ~33 TiB total, 9.0 TiB used, 24 TiB avail +- **Pools**: 13 pools, 533 PGs (532 active+clean, 1 scrubbing) ## OSD Layout | OSD | Class | Size | Host | Reweight | Notes | |-----|-------|------|------|----------|-------| -| 0 | ssd | 188 GB | proxmox2 | — | **DESTROYED 2026-07-25** (92% wear) | -| 1 | hdd | 3.7 TiB | n5pro | 1.0 | Large HDD | -| 2 | ssd | 233 GB | proxmox2 | — | **DESTROYED 2026-07-25** (slow ops, 81% full) | -| 3 | ssd | 238 GB | proxmox4 | 0.95 | | -| 4 | ssd | 233 GB | proxmox3 | 0.95 | | -| 5 | ssd | 238 GB | proxmox5 | 0.90 | 80% full | -| 6 | hdd | 2.8 TiB | ubuntu | 1.0 | Large HDD | -| 7 | hdd | 500 GB | proxmox7 | 1.0 | Was 0.80, reweighted 2026-07-24 | -| 8 | hdd | 2.8 TiB | ubuntu | 1.0 | BlueFS spillover | +| 1 | hdd | 3.7 TiB | n5pro | 1.0 | | +| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host | +| 3 | ssd | 233 GB | proxmox4 | 1.0 | | +| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down | +| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down | +| 6 | hdd | 3.6 TiB | n5pro | 1.0 | | +| 7 | hdd | 931 GB | proxmox7 | 0.95 | | +| 8 | hdd | 3.6 TiB | ubuntu | 1.0 | | | 9 | ssd | 1.9 TiB | n5pro | 1.0 | | -| 10 | hdd | 300 GB | proxmox6 | 1.0 | Very small HDD | -| 11 | hdd | 2.8 TiB | n5pro | 1.0 | Large HDD | +| 10 | hdd | 931 GB | proxmox6 | 0.95 | | +| 11 | hdd | 2.8 TiB | n5pro | 1.0 | | | 12 | ssd | 1.9 TiB | n5pro | 1.0 | | -| 13 | ssd | 1.8 TiB | ms-a2-1 | 1.0 | | -| 14 | ssd | 1.8 TiB | ms-a2-2 | 1.0 | New 2026-07-24, nvme0n1 | +| 13 | ssd | 1.8 TiB | ms-a2-1 | 0.95 | | +| 14 | ssd | 1.8 TiB | ms-a2-2 | 0.95 | | +| 15 | ssd | 1.8 TiB | ubuntu | 1.0 | New | +| 17 | hdd | 3.6 TiB | ubuntu | 1.0 | New | +| 18 | hdd | 3.6 TiB | ubuntu | 1.0 | New | + +> OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD. +> OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded). ## Pools @@ -44,7 +51,7 @@ modified: "2026-07-25" | 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on | | 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | | | 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on | -| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true (was 120, equalized to 112) | +| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true | | 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB | | 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites | | 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | | @@ -58,46 +65,22 @@ modified: "2026-07-25" ## Known Issues -### Weight Imbalance Causing Placement Failures (2026-07-24) -HDD hosts have extreme weight disparity: n5pro=10.15TB, ubuntu=5.49TB, proxmox7=0.50TB, proxmox6=0.30TB. -CRUSH host-level selection (rule 1) often picks only 2 of 4 HDD hosts → up sets with 2 OSDs instead of 3. -Result: PGs stuck in `active+clean+remapped` because up set < min_size. +### HEALTH_WARN: Insecure Key Types (2026-09-26) +Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4. +Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not already set). -**Mitigation (2026-07-24)**: -1. Reweighted osd.7 from 0.80 → 1.0 → fixed EC pool 8.3d (NONE → osd.7) -2. Equalized pool 6 pg_num 120 → 112 + nopgchange=true -3. Manual pg-upmap for stuck PGs: 5.13 → [1,6,7], 6.6c → [11,8,7], 6.58 → [11,6,7] -4. All `clean+remapped` eliminated. Triggered rebalancing wave (43 PGs backfilling at 26 MiB/s). +### Small SSDs causing reweightdown +osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity. +Consider removing from CRUSH or replacing with larger drives. -**Long-term**: Small HDDs (osd.7 0.5TB, osd.10 0.3TB) cause CRUSH placement failures. Replace with larger disks or create separate CRUSH root for large HDDs only. - -### Pool 6 pg_num/pgp_num Mismatch (Fixed 2026-07-24) -Pool hdd_disk had pg_num=120, pgp_num=112 (autoscaler reducing to 32). -Equalized pg_num to 112. Set nopgchange=true to prevent further autoscaler interference. - -### BlueFS Spillover on osd.8 -osd.8 spilled 128KiB metadata from db device (2.1GiB of 30GiB) to slow device. -Cosmetic warning, no data risk. Fix: `ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-8` - -### Slow Operations on osd.2 and osd.7 -osd.2 (81% full, fragmentation 0.80) and osd.7 (small HDD) experience slow BlueStore ops. -osd.2 NVMe has 92% wear — candidate for replacement. - -### osd.0 NVMe Wear -92% Wear, Critical Warning → Austausch planen. - -### EC Pool k=4+m=1 — No Rebalance Headroom -With 5 OSDs kein Rebalance Headroom. Siehe Solution Doc: `docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md` - -## RBD Management -- Proxmox RBD Double-Mount Deadlock Pitfall: Niemals `pct mount` und `pct exec` gleichzeitig auf demselben Container -- Siehe Solution Doc: `docs/solutions/bug-fixes/2026-07-23-proxmox-rbd-double-mount-deadlock.md` +### worker-04 (VM 139) NotReady in K8s +Node offline — not a Ceph issue but affects Ceph CSI attachments. ## Access -- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.92` +- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50` - Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs` -- Mon nodes: proxmox5, proxmox7, proxmox4 -- Mgr: proxmox5 (active) +- Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro +- Mgr: n5pro (active) ## Related Skills - `ceph-cluster-administration` (devops) diff --git a/systems/proxmox-cluster.md b/systems/proxmox-cluster.md index 86d4caa..2ed918d 100644 --- a/systems/proxmox-cluster.md +++ b/systems/proxmox-cluster.md @@ -3,14 +3,14 @@ title: Proxmox VE Cluster category: systems tags: [proxmox, virtualization, lxc, qemu, pve] created: "2026-04-28" -modified: "2026-07-24" +modified: "2026-09-26" --- # Proxmox VE Cluster ## Cluster-Konfiguration -- **Version:** PVE 9.2.10, Kernel 7.0.14-12-pve (upgraded 2026-08-17) -- **Nodes:** 9 (Quorum OK) +- **Version:** PVE 9.2.20, Kernel 7.0.14-19-pve (upgraded 2026-09-25) +- **Nodes:** 8 (Quorum OK, proxmox2 dauerhaft entfernt) - **Hypervisoren:** 10.0.20.x - **Guests:** ~30 LXC + ~10 QEMU VMs @@ -53,10 +53,10 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs - Worker-05 (VM 102) läuft auf ms-a2-2 mit funktionierendem GPU-Passthrough ## Bekannte Probleme -- osd.0 NVMe 92% Wear — Austausch planen -- osd.2/5 nearfull (93-94%) — entlasten -- CT110 kaputte libc — Reparatur ausstehend -- ms-a2-1 GPU-Passthrough: Config gefixt, Reboot zur Verifikation ausstehend +- CT110 kaputte libc — Reparatur ausstehend (still stopped) +- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen +- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation +- worker-04 (VM 139) NotReady im K8s Cluster — Node offline ## Related - [[systems/ceph-cluster]]