Auto-sync: 2026-09-29
This commit is contained in:
@@ -25,9 +25,13 @@ modified: "2026-09-17"
|
||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
||||
|
||||
## Storage
|
||||
- **Ceph CSI**: ceph-flash (fast), ceph-hdd (bulk)
|
||||
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
||||
- **⚠️ StorageClass IaC (2026-09-29):** `ceph-flash` + `ceph-hdd-replica` waren MANUELL erstellt (nicht in Git). Jetzt committed unter `clusters/main/storage/`. Alle RBD SCs MÜSSEN `controller-expand-secret-name/namespace` haben — fehlt dies, schlagen Volume Expansions fehl ("provided secret is empty") und triggern Retry-Loops die API Server überlasten.
|
||||
- **⚠️ PVs snapshotten SC-Parameter zur Provisionierungszeit** — SC fixen reicht NICHT; bestehende PVs brauchen individuellen Patch mit `controllerExpandSecretRef` falls sie ohne erstellt wurden.
|
||||
- **⚠️ StorageClass `parameters` sind IMMUTABLE** — können nicht gepatched werden, müssen gelöscht+neu erstellt werden.
|
||||
- **⚠️ Snapshot-CRD-Flavor (seit Rebuild 2026-08-01):** `volumesnapshotclasses`-CRD (RKE2-Addon `rke2-snapshot-controller-crd`) hat **FLAT-Schema** — `driver`/`deletionPolicy`/`parameters` auf TOP-LEVEL, `spec:` existiert nicht im Schema (Upstream-CRD wäre nested!). Niemals struktur-intuitiv "reparieren" — CRD-Schema lesen. Live-VSCs: `ceph-rbd-snapclass` (default) + `cephfs-snapclass`
|
||||
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`; Specs **100Gi data / 20Gi WAL** (grow-only — Shrink wird von CNPG-Admission verboten, Git immer nach oben alignieren)
|
||||
- **⚠️ etcd auf Ceph RBD (~20ms WAL fsync auf allen CP-Nodes)** — architektonisches Risiko; unter API Server Write Pressure → gRPC DeadlineExceeded Cascade. Lokales NVMe für etcd data dirs empfohlen.
|
||||
|
||||
## GitOps
|
||||
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
|
||||
|
||||
Reference in New Issue
Block a user