144 lines
6.1 KiB
Markdown
144 lines
6.1 KiB
Markdown
---
|
|
title: Proxmox VE Cluster
|
|
category: systems
|
|
tags: [proxmox, virtualization, lxc, qemu, pve]
|
|
created: "2026-04-28"
|
|
modified: "2026-09-30"
|
|
---
|
|
|
|
# Proxmox VE Cluster
|
|
|
|
## Cluster-Konfiguration
|
|
- **Version:** PVE 9.2.20, Kernel 7.0.14-19-pve (upgraded 2026-09-25)
|
|
- **Nodes:** 8 (Quorum OK, proxmox2 dauerhaft entfernt)
|
|
- **Hypervisoren:** 10.0.20.x
|
|
- **Guests:** ~30 LXC + ~10 QEMU VMs
|
|
|
|
## Storage
|
|
- **vm_disks:** Primärer Storage für alle VMs/CTs (SSD)
|
|
- **hdd_templates:** CT Templates
|
|
- **Ceph RBD:** ceph-flash, ceph-hdd Pools (über K8s CSI)
|
|
|
|
## Netzwerk
|
|
- VLAN-basiert, Bridge vmbr0
|
|
- IP-Schema: 10.0.X.Y — siehe [[concepts/network-architecture]]
|
|
|
|
## Fluent Bit (Logging)
|
|
- Alle 9 Hosts haben Fluent Bit aktiv
|
|
- Inputs: systemd journal (pve*, corosync, pacemaker, ceph, zfs, smartd), auth.log, pveproxy/access.log, pvedaemon.log, cluster.log
|
|
- **PVE Tasks** (`/var/log/pve/tasks/index`): UPID-Format (Node, PID, Task-Type, VMID, User, Status)
|
|
- **Ceph Audit** (`/var/log/ceph/ceph.audit.log`): OSD/Pool/RBD Operationen
|
|
- Output → Loki (10.0.30.207:3100)
|
|
- Siehe [[systems/loki-fluentbit]]
|
|
|
|
## Wichtige Befehle
|
|
```bash
|
|
pvecm status # Cluster-Quorum
|
|
pct status <vmid> # Container-Status
|
|
pct start/stop <vmid> # Container starten/stoppen
|
|
qm status <vmid> # VM-Status
|
|
pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
|
```
|
|
|
|
## SSH-Zugriff
|
|
- Key: `id_ed25519_proxmox` (funktioniert für 10.0.20.x Hosts)
|
|
- Siehe [[reference/ssh-keys]]
|
|
|
|
## GPU Passthrough (AMD 1002:13c0)
|
|
- **ms-a2-1** (10.0.20.92): VFIO config gefixt 2026-07-24 — siehe [Solution Doc](../../../docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md)
|
|
- **ms-a2-2** (10.0.20.93): Funktioniert seit Initialisierung
|
|
- Benötigte modprobe.d Config:
|
|
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
|
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
|
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
|
|
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
|
|
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
|
|
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
|
|
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
|
|
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
|
|
|
|
## Bekannte Probleme
|
|
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
|
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
|
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
|
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
|
|
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
|
|
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
|
|
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
|
|
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
|
|
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
|
|
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
|
|
|
|
## HA Rules (PVE 9.2 Rules System)
|
|
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
|
Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
|
|
|
**Resource-Affinity (negative = anti-collocation):**
|
|
|
|
*RKE2 Control Plane (etcd-Quorum braucht 2/3):*
|
|
- `rke2-cp-anti-112-122`: vm:112 ↔ vm:122 — nie auf gleichem Host
|
|
- `rke2-cp-anti-112-126`: vm:112 ↔ vm:126 — nie auf gleichem Host
|
|
- `rke2-cp-anti-122-126`: vm:122 ↔ vm:126 — nie auf gleichem Host
|
|
|
|
*RKE2 Worker:*
|
|
- `rke2-worker-anti-102-128`: vm:102 ↔ vm:128 — nie auf gleichem Host
|
|
- `rke2-worker-anti-102-139`: vm:102 ↔ vm:139 — nie auf gleichem Host
|
|
- `rke2-worker-anti-128-139`: vm:128 ↔ vm:139 — nie auf gleichem Host
|
|
|
|
*Galera:*
|
|
- `galera-anti-300-301`: vm:300 ↔ vm:301 — nie auf gleichem Host
|
|
- `galera-anti-300-302`: vm:300 ↔ vm:302 — nie auf gleichem Host
|
|
- `galera-anti-301-302`: vm:301 ↔ vm:302 — nie auf gleichem Host
|
|
|
|
*MaxScale:*
|
|
- `maxscale-anti-310-311`: vm:310 ↔ vm:311 — nie auf gleichem Host
|
|
|
|
**Node-Affinity (non-strict, failover allowed):**
|
|
|
|
*RKE2 CP:*
|
|
- `na-vm112`: vm:112 → proxmox3, proxmox5, proxmox7
|
|
- `na-vm122`: vm:122 → ms-a2-1, ms-a2-2
|
|
- `na-vm126`: vm:126 → proxmox4, proxmox5, proxmox6
|
|
|
|
*RKE2 Worker:*
|
|
- `na-vm102`: vm:102 → proxmox5, proxmox6, proxmox7 (NICHT ms-a2-2!)
|
|
- `na-vm128`: vm:128 → ms-a2-2, proxmox5, proxmox7
|
|
- `na-vm139`: vm:139 → n5pro, proxmox3, proxmox4
|
|
|
|
*Hermes:*
|
|
- `na-vm230`: vm:230 → n5pro, proxmox5, proxmox6
|
|
|
|
*Galera/MaxScale:*
|
|
- `na-vm300`: vm:300 → n5pro, proxmox3, proxmox4
|
|
- `na-vm301`: vm:301 → proxmox6, proxmox5, proxmox7
|
|
- `na-vm302`: vm:302 → ms-a2-2, ms-a2-1
|
|
- `na-vm310`: vm:310 → proxmox7, proxmox5, proxmox4
|
|
- `na-vm311`: vm:311 → ms-a2-1, ms-a2-2
|
|
|
|
**Aktuelle Verteilung (alle Anti-Collocation erfüllt):**
|
|
| Role | VM | Node |
|
|
|------|----|------|
|
|
| RKE2 CP-01 | 112 | proxmox3 |
|
|
| RKE2 CP-02 | 122 | ms-a2-1 |
|
|
| RKE2 CP-03 | 126 | proxmox4 |
|
|
| RKE2 Worker-01 | 128 | proxmox5 |
|
|
| RKE2 Worker-04 | 139 | n5pro |
|
|
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
|
| Hermes-Agent-01 | 230 | n5pro |
|
|
| Galera db1 | 300 | n5pro |
|
|
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
|
|
| Galera db3 | 302 | ms-a2-2 |
|
|
| MaxScale-01 | 310 | proxmox7 |
|
|
| MaxScale-02 | 311 | ms-a2-1 |
|
|
|
|
> ⚠️ PVE 9.2 Constraints:
|
|
> - Resources in resource-affinity rules dürfen keine multi-priority node-affinity haben (gleiche Priorität für alle Nodes erforderlich).
|
|
> - `ha-manager add` MUSS vor `ha-manager rules add` kommen — sonst "cannot use unmanaged resource".
|
|
> - Bei gleichzeitigem HA-Add + Anti-Collocation-Violation kann HA-Manager deadlocks (beide VMs auf `migrate` fest). Lösung: eine VM temporär aus HA entfernen, manuell migrieren, dann re-add.
|
|
> - Online-Migration von VMs mit hohen Memory-Writes (>12GB dirty pages) kann `broken pipe` fehlschlagen. Offline-Migration (stop→migrate→start) als Fallback.
|
|
|
|
## Related
|
|
- [[systems/ceph-cluster]]
|
|
- [[systems/rke2-kubernetes]]
|
|
- [[reference/ip-map]]
|