Auto-sync: 2026-09-26

This commit is contained in:
Dominik Schön
2026-09-26 22:00:22 +00:00
parent f3039765ed
commit 02f6a8e2b0
5 changed files with 68 additions and 83 deletions
+37 -54
View File
@@ -3,37 +3,44 @@ title: Ceph Cluster
category: systems
tags: [ceph, storage, rbd, ec-pool, osd]
created: "2026-07-24"
modified: "2026-07-25"
modified: "2026-09-26"
---
# Ceph Cluster
## Overview
- **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
- **Health**: HEALTH_OK (recovery complete after OSD 0+2 drain, 0.3% misplaced settling)
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH 2026-07-25), 3 MONs (proxmox5/7/4), MGR on proxmox5
- **OSDs**: 13 (8 SSD, 5 HDD), all up/in — OSDs 0+2 destroyed+purged 2026-07-25
- **Capacity**: ~22 TiB total, 6.0 TiB used
- **Health**: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed)
- **Version**: 20.2.4 (tentacle) — all 17 OSDs
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1)
- **OSDs**: 17 (10 HDD, 7 SSD), all up/in
- **Capacity**: ~33 TiB total, 9.0 TiB used, 24 TiB avail
- **Pools**: 13 pools, 533 PGs (532 active+clean, 1 scrubbing)
## OSD Layout
| OSD | Class | Size | Host | Reweight | Notes |
|-----|-------|------|------|----------|-------|
| 0 | ssd | 188 GB | proxmox2 | — | **DESTROYED 2026-07-25** (92% wear) |
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | Large HDD |
| 2 | ssd | 233 GB | proxmox2 | — | **DESTROYED 2026-07-25** (slow ops, 81% full) |
| 3 | ssd | 238 GB | proxmox4 | 0.95 | |
| 4 | ssd | 233 GB | proxmox3 | 0.95 | |
| 5 | ssd | 238 GB | proxmox5 | 0.90 | 80% full |
| 6 | hdd | 2.8 TiB | ubuntu | 1.0 | Large HDD |
| 7 | hdd | 500 GB | proxmox7 | 1.0 | Was 0.80, reweighted 2026-07-24 |
| 8 | hdd | 2.8 TiB | ubuntu | 1.0 | BlueFS spillover |
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
| 3 | ssd | 233 GB | proxmox4 | 1.0 | |
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
| 7 | hdd | 931 GB | proxmox7 | 0.95 | |
| 8 | hdd | 3.6 TiB | ubuntu | 1.0 | |
| 9 | ssd | 1.9 TiB | n5pro | 1.0 | |
| 10 | hdd | 300 GB | proxmox6 | 1.0 | Very small HDD |
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | Large HDD |
| 10 | hdd | 931 GB | proxmox6 | 0.95 | |
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | |
| 12 | ssd | 1.9 TiB | n5pro | 1.0 | |
| 13 | ssd | 1.8 TiB | ms-a2-1 | 1.0 | |
| 14 | ssd | 1.8 TiB | ms-a2-2 | 1.0 | New 2026-07-24, nvme0n1 |
| 13 | ssd | 1.8 TiB | ms-a2-1 | 0.95 | |
| 14 | ssd | 1.8 TiB | ms-a2-2 | 0.95 | |
| 15 | ssd | 1.8 TiB | ubuntu | 1.0 | New |
| 17 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
| 18 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
> OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD.
> OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded).
## Pools
@@ -44,7 +51,7 @@ modified: "2026-07-25"
| 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on |
| 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | |
| 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on |
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true (was 120, equalized to 112) |
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true |
| 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB |
| 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites |
| 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | |
@@ -58,46 +65,22 @@ modified: "2026-07-25"
## Known Issues
### Weight Imbalance Causing Placement Failures (2026-07-24)
HDD hosts have extreme weight disparity: n5pro=10.15TB, ubuntu=5.49TB, proxmox7=0.50TB, proxmox6=0.30TB.
CRUSH host-level selection (rule 1) often picks only 2 of 4 HDD hosts → up sets with 2 OSDs instead of 3.
Result: PGs stuck in `active+clean+remapped` because up set < min_size.
### HEALTH_WARN: Insecure Key Types (2026-09-26)
Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4.
Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not already set).
**Mitigation (2026-07-24)**:
1. Reweighted osd.7 from 0.80 → 1.0 → fixed EC pool 8.3d (NONE → osd.7)
2. Equalized pool 6 pg_num 120 → 112 + nopgchange=true
3. Manual pg-upmap for stuck PGs: 5.13 → [1,6,7], 6.6c → [11,8,7], 6.58 → [11,6,7]
4. All `clean+remapped` eliminated. Triggered rebalancing wave (43 PGs backfilling at 26 MiB/s).
### Small SSDs causing reweightdown
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
Consider removing from CRUSH or replacing with larger drives.
**Long-term**: Small HDDs (osd.7 0.5TB, osd.10 0.3TB) cause CRUSH placement failures. Replace with larger disks or create separate CRUSH root for large HDDs only.
### Pool 6 pg_num/pgp_num Mismatch (Fixed 2026-07-24)
Pool hdd_disk had pg_num=120, pgp_num=112 (autoscaler reducing to 32).
Equalized pg_num to 112. Set nopgchange=true to prevent further autoscaler interference.
### BlueFS Spillover on osd.8
osd.8 spilled 128KiB metadata from db device (2.1GiB of 30GiB) to slow device.
Cosmetic warning, no data risk. Fix: `ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-8`
### Slow Operations on osd.2 and osd.7
osd.2 (81% full, fragmentation 0.80) and osd.7 (small HDD) experience slow BlueStore ops.
osd.2 NVMe has 92% wear — candidate for replacement.
### osd.0 NVMe Wear
92% Wear, Critical Warning → Austausch planen.
### EC Pool k=4+m=1 — No Rebalance Headroom
With 5 OSDs kein Rebalance Headroom. Siehe Solution Doc: `docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`
## RBD Management
- Proxmox RBD Double-Mount Deadlock Pitfall: Niemals `pct mount` und `pct exec` gleichzeitig auf demselben Container
- Siehe Solution Doc: `docs/solutions/bug-fixes/2026-07-23-proxmox-rbd-double-mount-deadlock.md`
### worker-04 (VM 139) NotReady in K8s
Node offline — not a Ceph issue but affects Ceph CSI attachments.
## Access
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.92`
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50`
- Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs`
- Mon nodes: proxmox5, proxmox7, proxmox4
- Mgr: proxmox5 (active)
- Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro
- Mgr: n5pro (active)
## Related Skills
- `ceph-cluster-administration` (devops)