Files

103 lines
4.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: Ceph Cluster
category: systems
tags: [ceph, storage, rbd, ec-pool, osd]
created: "2026-07-24"
modified: "2026-09-26"
---
# Ceph Cluster
## Overview
- **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
- **Health**: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed)
- **Version**: 20.2.4 (tentacle) — all 17 OSDs
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1)
- **OSDs**: 17 (10 HDD, 7 SSD), all up/in
- **Capacity**: ~33 TiB total, 9.0 TiB used, 24 TiB avail
- **Pools**: 13 pools, 533 PGs (532 active+clean, 1 scrubbing)
## OSD Layout
| OSD | Class | Size | Host | Reweight | Notes |
|-----|-------|------|------|----------|-------|
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
| 7 | hdd | 931 GB | proxmox7 | 0.95 | |
| 8 | hdd | 3.6 TiB | ubuntu | 1.0 | |
| 9 | ssd | 1.9 TiB | n5pro | 1.0 | |
| 10 | hdd | 931 GB | proxmox6 | 0.95 | |
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | |
| 12 | ssd | 1.9 TiB | n5pro | 1.0 | |
| 13 | ssd | 1.8 TiB | ms-a2-1 | 0.95 | |
| 14 | ssd | 1.8 TiB | ms-a2-2 | 0.95 | |
| 15 | ssd | 1.8 TiB | ubuntu | 1.0 | New |
| 17 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
| 18 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
> OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD.
> OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded).
## Pools
| Pool | Name | Type | Size | Min | CRUSH Rule | PGs | Notes |
|------|------|------|------|-----|------------|-----|-------|
| 1 | cephfs_data | replicated | 3 | 2 | 0 (any) | 32 | autoscale off |
| 2 | cephfs_metadata | replicated | 3 | 2 | 2 (ssd) | 32 | autoscale off |
| 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on |
| 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | |
| 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on |
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true |
| 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB |
| 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites |
| 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | |
| 10 | .rgw.root | replicated | 3 | 2 | 0 (any) | 1 | |
## CRUSH Rules
- **Rule 0** (replicated_rule): default root, host-level placement
- **Rule 1** (replicated_hdd): default~hdd, host-level placement
- **Rule 2** (replicated_ssd): default~ssd, host-level placement
- **Rule 3** (media_ec): default~hdd, OSD-level placement (choose_indep)
## Known Issues
### HEALTH_WARN: Insecure Key Types (2026-09-26)
Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4.
Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not already set).
### Small SSDs causing reweightdown
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
Consider removing from CRUSH or replacing with larger drives.
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
thermal throttle T1 3×.**
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
### worker-04 (VM 139) NotReady in K8s
Node offline — not a Ceph issue but affects Ceph CSI attachments.
## Access
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50`
- Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs`
- Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro
- Mgr: n5pro (active)
## Related Skills
- `ceph-cluster-administration` (devops)
## Related
- [[systems/proxmox-cluster]]
- [[systems/rke2-kubernetes]] (Ceph CSI)