Auto-sync: 2026-09-30
This commit is contained in:
@@ -1,5 +1,14 @@
|
||||
# Memory Log
|
||||
|
||||
## [2026-09-30] k8s-cp01-relief-descheduler | Workload-Migration + Descheduler-Deployment
|
||||
- cp-01 bei 93% RAM → hindsight-api + hindsight-postgres + paperless nach worker-05 migriert (cordon/delete/uncordon). cp-01 RAM 93%→43%.
|
||||
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren bis konvergiert.
|
||||
- hindsight-api Live-Pin (Sep-17-Hotfix) war nie in Git → Live-Patch nötig.
|
||||
- Descheduler v0.30.0 deployt (raw Manifests, nicht Helm — Chart hat params/args-Bug + falsche API-Version). Policy: LowNodeUtilization (30/50/20 → 70/70/60), 2m Intervall, nodeFit auf Profile-Ebene.
|
||||
- RBAC-Lektion: Descheduler braucht `list/watch` auf `namespaces` in ClusterRole.
|
||||
- ArgoCD repoURL: `ssh://git@10.0.30.200:22/...` (mit `git@` User-Prefix, wie alle anderen Apps).
|
||||
- Solution Doc + INDEX + Wiki rke2-kubernetes.md aktualisiert. Commits b955d7e→8c1ecff.
|
||||
|
||||
## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur
|
||||
- Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen.
|
||||
- Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate.
|
||||
|
||||
Reference in New Issue
Block a user