Files
memory/systems/proxmox-cluster.md
T
2026-09-26 22:00:22 +00:00

2.3 KiB

title, category, tags, created, modified
title category tags created modified
Proxmox VE Cluster systems
proxmox
virtualization
lxc
qemu
pve
2026-04-28 2026-09-26

Proxmox VE Cluster

Cluster-Konfiguration

  • Version: PVE 9.2.20, Kernel 7.0.14-19-pve (upgraded 2026-09-25)
  • Nodes: 8 (Quorum OK, proxmox2 dauerhaft entfernt)
  • Hypervisoren: 10.0.20.x
  • Guests: ~30 LXC + ~10 QEMU VMs

Storage

  • vm_disks: Primärer Storage für alle VMs/CTs (SSD)
  • hdd_templates: CT Templates
  • Ceph RBD: ceph-flash, ceph-hdd Pools (über K8s CSI)

Netzwerk

Fluent Bit (Logging)

  • Alle 9 Hosts haben Fluent Bit aktiv
  • Inputs: systemd journal (pve*, corosync, pacemaker, ceph, zfs, smartd), auth.log, pveproxy/access.log, pvedaemon.log, cluster.log
  • PVE Tasks (/var/log/pve/tasks/index): UPID-Format (Node, PID, Task-Type, VMID, User, Status)
  • Ceph Audit (/var/log/ceph/ceph.audit.log): OSD/Pool/RBD Operationen
  • Output → Loki (10.0.30.207:3100)
  • Siehe systems/loki-fluentbit

Wichtige Befehle

pvecm status          # Cluster-Quorum
pct status <vmid>     # Container-Status
pct start/stop <vmid> # Container starten/stoppen
qm status <vmid>      # VM-Status
pvesh get /cluster/resources --type vm  # Alle VMs/CTs

SSH-Zugriff

GPU Passthrough (AMD 1002:13c0)

  • ms-a2-1 (10.0.20.92): VFIO config gefixt 2026-07-24 — siehe Solution Doc
  • ms-a2-2 (10.0.20.93): Funktioniert seit Initialisierung
  • Benötigte modprobe.d Config:
    • blacklist amdgpu + blacklist drm + blacklist drm_kms_helper
    • options vfio-pci ids=1002:13c0 + softdep amdgpu pre: vfio-pci
  • Worker-05 (VM 102) läuft auf ms-a2-2 mit funktionierendem GPU-Passthrough

Bekannte Probleme

  • CT110 kaputte libc — Reparatur ausstehend (still stopped)
  • osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
  • osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
  • worker-04 (VM 139) NotReady im K8s Cluster — Node offline