Files
memory/systems/ceph-cluster.md
T

4.5 KiB
Raw Blame History

title, category, tags, created, modified
title category tags created modified
Ceph Cluster systems
ceph
storage
rbd
ec-pool
osd
2026-07-24 2026-09-26

Ceph Cluster

Overview

  • Cluster ID: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
  • Health: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed)
  • Version: 20.2.4 (tentacle) — all 17 OSDs
  • Nodes: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1)
  • OSDs: 17 (10 HDD, 7 SSD), all up/in
  • Capacity: ~33 TiB total, 9.0 TiB used, 24 TiB avail
  • Pools: 13 pools, 533 PGs (532 active+clean, 1 scrubbing)

OSD Layout

OSD Class Size Host Reweight Notes
1 hdd 3.7 TiB n5pro 1.0
2 ssd 1.8 TiB ubuntu 1.0 Moved to ubuntu host
3 ssd 233 GB proxmox4 0.05 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull
4 ssd 227 GB proxmox3 0.30 Small, reweighted down
5 ssd 150 GB proxmox5 0.30 Small, reweighted down
6 hdd 3.6 TiB n5pro 1.0
7 hdd 931 GB proxmox7 0.95
8 hdd 3.6 TiB ubuntu 1.0
9 ssd 1.9 TiB n5pro 1.0
10 hdd 931 GB proxmox6 0.95
11 hdd 2.8 TiB n5pro 1.0
12 ssd 1.9 TiB n5pro 1.0
13 ssd 1.8 TiB ms-a2-1 0.95
14 ssd 1.8 TiB ms-a2-2 0.95
15 ssd 1.8 TiB ubuntu 1.0 New
17 hdd 3.6 TiB ubuntu 1.0 New
18 hdd 3.6 TiB ubuntu 1.0 New

OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD. OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded).

Pools

Pool Name Type Size Min CRUSH Rule PGs Notes
1 cephfs_data replicated 3 2 0 (any) 32 autoscale off
2 cephfs_metadata replicated 3 2 2 (ssd) 32 autoscale off
3 vm_disks replicated 3 2 2 (ssd) 128 autoscale on
4 .mgr replicated 3 2 2 (ssd) 1
5 rbd replicated 3 2 1 (hdd) 32 autoscale on
6 hdd_disk replicated 3 2 1 (hdd) 112 nopgchange=true
7 tm_disks replicated 2 2 1 (hdd) 128 target_size 2TiB
8 media_ec erasure 4+1 5 4 3 (hdd, osd-level) 128 ec_overwrites
9 media_meta replicated 3 2 0 (any) 32
10 .rgw.root replicated 3 2 0 (any) 1

CRUSH Rules

  • Rule 0 (replicated_rule): default root, host-level placement
  • Rule 1 (replicated_hdd): default~hdd, host-level placement
  • Rule 2 (replicated_ssd): default~ssd, host-level placement
  • Rule 3 (media_ec): default~hdd, OSD-level placement (choose_indep)

Known Issues

HEALTH_WARN: Insecure Key Types (2026-09-26)

Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4. Fix: ceph config set mon mon_allow_insecure_global_id_reclaim false (if not already set).

Small SSDs causing reweightdown

osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity. Consider removing from CRUSH or replacing with larger drives.

NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)

Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out. Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device. Details: patterns/ceph-dead-nvme-resurrection (PAT-014). Update 2026-09-30 (Abend): SMART- Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C, thermal throttle T1 3×.

ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)

Ohne DNS-Record wirft die PVE-GUI hostname lookup 'ubuntu' failed (500). Fix: hosts-Eintrag 10.0.20.100 ubuntu fleetweit auf allen 8 Nodes.

worker-04 (VM 139) NotReady in K8s

Node offline — not a Ceph issue but affects Ceph CSI attachments.

Access

  • SSH to Proxmox hosts: ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50
  • Ceph commands: ceph status, ceph osd tree, ceph pg dump pgs
  • Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro
  • Mgr: n5pro (active)
  • ceph-cluster-administration (devops)