Auto-sync: 2026-09-29

This commit is contained in:
Dominik Schön
2026-09-29 22:00:35 +00:00
parent f09c6ef414
commit d487b596d7
5 changed files with 84 additions and 9 deletions
+17
View File
@@ -1,5 +1,14 @@
# Memory Log
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
- Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min)
- Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05
- Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2)
- Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes)
- Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
@@ -213,3 +222,11 @@
- Created 2 concept pages: gitops-workflow, credential-policy
- Rewrote index.md with new structure + Memory Layer Architecture table
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
## [2026-09-29] update | VM302 Zombie-Recovery nach Migration
- VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot
- Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause
- Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min
- Endstand: 3/3 Synced, Primary, MaxScale alle Server Running
- Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md
- Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)