Compare commits
23
Commits
c96ad03ed7
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b527e4ff3c | ||
|
|
3a57db957e | ||
|
|
cc2f237ca3 | ||
|
|
531d6bfbdd | ||
|
|
a68665eca0 | ||
|
|
d487b596d7 | ||
|
|
f09c6ef414 | ||
|
|
dfd3add11b | ||
|
|
d1464eb9e3 | ||
|
|
02f6a8e2b0 | ||
|
|
f3039765ed | ||
|
|
1d036dbcdc | ||
|
|
fb47c2a182 | ||
|
|
b2e9512a0e | ||
|
|
7c36a7600a | ||
|
|
23e3702f0a | ||
|
|
5cf6a573f1 | ||
|
|
bbbe4f8985 | ||
|
|
6a7aed6f48 | ||
|
|
5334f8ec07 | ||
|
|
971edfea8e | ||
|
|
222f286adc | ||
|
|
c392fc45f0 |
@@ -14,7 +14,7 @@ Infra-Changes **bevorzugt** über GitOps. Geht nicht immer. Ausnahmen müssen vo
|
|||||||
## Workflow
|
## Workflow
|
||||||
1. IaC Repo lokal auschecken (`/root/iac-homelab` auf VM200)
|
1. IaC Repo lokal auschecken (`/root/iac-homelab` auf VM200)
|
||||||
2. Ändern (Tofu/Ansible/K8s Manifeste)
|
2. Ändern (Tofu/Ansible/K8s Manifeste)
|
||||||
3. Commit + Push (zu BEIDEN Gitea Remotes: origin + k8s)
|
3. Commit + Push zum EINEN kanonischen Remote: `ssh://git@10.0.30.202:22/dominik/iac-homelab.git` (Dual-Push-Habit führte 09/26 zur Ghost-Regression über CT108 — NIEMALS wieder zweites Remote pflegen)
|
||||||
4. ArgoCD sync (oder auto-sync)
|
4. ArgoCD sync (oder auto-sync)
|
||||||
5. Verify (kubectl get, curl, etc.)
|
5. Verify (kubectl get, curl, etc.)
|
||||||
6. Lokale Kopie löschen
|
6. Lokale Kopie löschen
|
||||||
|
|||||||
@@ -0,0 +1,135 @@
|
|||||||
|
---
|
||||||
|
title: "Memory Layer Architecture"
|
||||||
|
category: concepts
|
||||||
|
tags: [memory, architecture, hindsight, lcm, wiki, separation]
|
||||||
|
created: "2026-09-27"
|
||||||
|
modified: "2026-09-27"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Memory Layer Architecture — Practical Rules
|
||||||
|
|
||||||
|
## Übersicht
|
||||||
|
|
||||||
|
7 Schichten mit klaren, nicht-überlappenden Rollen.
|
||||||
|
|
||||||
|
| # | Layer | System | Pfad/Ort | Rolle | Wann nutzen |
|
||||||
|
|---|-------|--------|----------|-------|-------------|
|
||||||
|
| L0 | Hot | MEMORY.md + USER.md | `~/.hermes/memories/` | System-Prompt Injection, komprimierte Facts | Jede Session (automatisch) |
|
||||||
|
| L1 | Curated | LLM Wiki | `~/.hermes/memory/` | Browsbares Wissen, git-synced | Wenn Details gebraucht werden |
|
||||||
|
| L2 | Semantic | Hindsight | K8s PostgreSQL | Vektor-Suche, semantisches Retrieval | Cross-Session Lookup |
|
||||||
|
| L3 | Procedural | Skills | `~/.hermes/skills/` | Wie-geht-es Prozeduren mit Pitfalls | Bei wiederkehrenden Tasks |
|
||||||
|
| L4 | Episodic | LCM + session_search | SQLite lcm.db / state.db | Rohe Gesprächsverläufe, Compaction | Innerhalb aktiver Session |
|
||||||
|
| L5 | Solutions | Solution Docs | `~/docs/solutions/` | Problem-Lösungs-Paare | Nach substanziellen Tasks |
|
||||||
|
| L6 | Project | AGENTS.md | pro Repo | Repo-spezifische Konventionen | Beim Arbeiten in einem Repo |
|
||||||
|
|
||||||
|
## Was wohin gehört — Entscheidungsbaum
|
||||||
|
|
||||||
|
```
|
||||||
|
Ist es ein Fact über Dominik persönlich?
|
||||||
|
→ USER.md (L0)
|
||||||
|
Ist es sein Kommunikationsstil, Safety-Rule, oder Workflow-Präferenz?
|
||||||
|
→ USER.md
|
||||||
|
Ist es ein persönlicher Umstand (Größe, Gewicht, Arbeitgeber)?
|
||||||
|
→ Hindsight (L2) — NICHT USER.md
|
||||||
|
|
||||||
|
Ist es ein Infra-Fakt (IP, Version, Konfiguration)?
|
||||||
|
→ LLM Wiki (L1) — systems/ oder reference/
|
||||||
|
Braucht er jede Session sofort?
|
||||||
|
→ MEMORY.md (L0) als 1-Zeiliger Pointer "→Wiki systems/xyz"
|
||||||
|
|
||||||
|
Ist es ein "Wie mache ich X"-Prozess?
|
||||||
|
→ Skills (L3)
|
||||||
|
|
||||||
|
Ist es ein "Wir hatten Problem Y, Lösung war Z"-Eintrag?
|
||||||
|
→ Solution Doc (L5) + Hindsight-Index (L2)
|
||||||
|
|
||||||
|
Ist es ein vergangenes Gesprächsereignis?
|
||||||
|
→ LCM (L4) kümmert sich automatisch darum
|
||||||
|
```
|
||||||
|
|
||||||
|
## L0 — MEMORY.md vs USER.md
|
||||||
|
|
||||||
|
### USER.md (1.375 chars max)
|
||||||
|
**Nur:** Dominiks Persönliche Präferenzen, Safety-Rules, Kommunikationsstil
|
||||||
|
- "NIEMALS ohne Approval löschen"
|
||||||
|
- "Action-first, kein Dom"
|
||||||
|
- "Crash+Alert statt Silent Skip"
|
||||||
|
- "Laya in Sessions nutzen"
|
||||||
|
|
||||||
|
**NICHT:** Infra-Fakten, IP-Adressen, Versionsnummern, Team-Struktur
|
||||||
|
|
||||||
|
### MEMORY.md (2.200 chars max)
|
||||||
|
**Nur:** Komprimierte Pointer auf Wiki-Seiten + nicht-wikiffähige Quick-Facts
|
||||||
|
- `Frigate v0.18 CT151. →Wiki systems/frigate`
|
||||||
|
- `Galera:VM300/301/302,VIP .70. →Wiki systems/galera-maxscale`
|
||||||
|
- Baufinanz-Zahlen (persönlich, nicht im Wiki)
|
||||||
|
- LinkedIn Persona (Workflow, nicht im Wiki)
|
||||||
|
|
||||||
|
**NICHT:** Vollständige Specs, Konfigurationsdetails, lange Beschreibungen
|
||||||
|
|
||||||
|
## L1 — LLM Wiki
|
||||||
|
|
||||||
|
### Wann erstellen/aktualisieren?
|
||||||
|
- Bei jeder infrastrukturellen Änderung (Compound-Learning Cycle Phase 4.7)
|
||||||
|
- Neue Seite wenn: neues System, neuer Service, neue architektonische Entscheidung
|
||||||
|
- Patch wenn: Versionswechsel, IP-Änderung, Known Issue hinzugekommen
|
||||||
|
|
||||||
|
### Wann NICHT?
|
||||||
|
- Procedures → Skills (L3)
|
||||||
|
- Problem-Solution Pairs → Solution Docs (L5)
|
||||||
|
- Credentials → 1Password
|
||||||
|
|
||||||
|
## L2 — Hindsight
|
||||||
|
|
||||||
|
### Was gehört rein?
|
||||||
|
- **Semantische Pointer** auf Solution Docs und wichtige Meilensteine
|
||||||
|
- **Cross-Session Kontext** der nicht im Wiki steht (z.B. "wir haben am Datum X entschieden...")
|
||||||
|
- **Episodisches Wissen** über Projekte, Entscheidungen, Learnings
|
||||||
|
|
||||||
|
### Was NICHT rein gehört?
|
||||||
|
- ❌ Infra-Topologie (→ L1 Wiki)
|
||||||
|
- ❌ Aktuelle IPs/Versionsnummern (→ L1 Wiki, werden schnell stale)
|
||||||
|
- ❌ Procedures (→ L3 Skills)
|
||||||
|
- ❌ Credentials (→ 1Password)
|
||||||
|
- ❌ Session-Fortschritt (→ L4 LCM)
|
||||||
|
|
||||||
|
### Bekanntes Problem: Hindsight Accumulation
|
||||||
|
Hindsight hat über Monate Infra-Fakten angesammelt die inzwischen teilweise
|
||||||
|
veraltet sind (9 vs 8 Nodes, alte OSD-Counts). Hindsight bietet keine
|
||||||
|
Delete-Funktion via Tools. Strategie:
|
||||||
|
1. Zukünftig nur noch semantische Pointer + Entscheidungen speichern
|
||||||
|
2. Topologie-Fakten nicht mehr via `hindsight_retain` ablegen
|
||||||
|
3. Veraltete Einträge ignorieren — Hindsight Ranking bevorzugt neuere Memories
|
||||||
|
|
||||||
|
## L4 — LCM (Local Conversation Memory)
|
||||||
|
|
||||||
|
### Rolle
|
||||||
|
- **Innerhalb aktiver Session:** Compaction, Summary-DAG, Fresh Tail
|
||||||
|
- **Cross-Session (lcm.db):** Rohe Nachrichten + Summary Nodes aller Sessions
|
||||||
|
- **Kein manuelles Management nötig** — läuft automatisch
|
||||||
|
|
||||||
|
### Wann manuell eingreifen?
|
||||||
|
- `lcm_status` bei Verdacht auf Compression-Problemen
|
||||||
|
- `lcm_grep` um exakte Phrasen in vergangenen Sessions zu finden
|
||||||
|
- `lcm_recall` für bedeutungsbasierte Cross-Session-Suche
|
||||||
|
- Niemals manuell Daten in LCM schreiben — es ist Auto-Managed
|
||||||
|
|
||||||
|
## L5 — Solution Docs
|
||||||
|
|
||||||
|
### Wann erstellen?
|
||||||
|
- Nach substanziellem Bugfix (≥5 Tool-Calls)
|
||||||
|
- Nach Architektur-Änderung
|
||||||
|
- Nach komplexer Migration
|
||||||
|
- Nach schwierigem Troubleshooting
|
||||||
|
|
||||||
|
### Format
|
||||||
|
`~/docs/solutions/{type}/{YYYY-MM-DD}-{slug}.md`
|
||||||
|
Types: `architecture/`, `bugfix/`, `migration/`, `workflow/`
|
||||||
|
|
||||||
|
Immer: Hindsight-Index (`hindsight_retain`) nach Erstellung
|
||||||
|
|
||||||
|
## Related
|
||||||
|
- [[systems/hindsight]] — Hindsight Setup, API
|
||||||
|
- [[systems/laya]] — Laya Classifier
|
||||||
|
- Skill: `memory-sync` — Wiki Maintenance Protocol
|
||||||
|
- Skill: `software-development/compound-learning` — Phase 4.7 Wiki Update
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
# User Drift Report — 2026-09-28
|
||||||
|
|
||||||
|
**Recent window:** Last 14 days (1884 msgs)
|
||||||
|
**Baseline:** Previous 90 days (6185 msgs)
|
||||||
|
|
||||||
|
## Detected Drifts
|
||||||
|
|
||||||
|
- **message_length**: Messages 98% longer (baseline: 1192 chars → recent: 2361)
|
||||||
|
- **new_focus**: New dominant topics: security
|
||||||
|
- **declining_focus**: Topics fading from focus: gitops
|
||||||
|
|
||||||
|
## Signal Summary
|
||||||
|
|
||||||
|
| Metric | Baseline | Recent |
|
||||||
|
|---|---|---|
|
||||||
|
| Avg msg length | 1192 | 2361 |
|
||||||
|
| Frustration rate | 4.2% | 6.6% |
|
||||||
|
| Correction rate | 7.1% | 8.8% |
|
||||||
|
| Action-first | 1.5% | 0.5% |
|
||||||
|
| Top topics | ceph, proxmox, k8s | security, backup, ceph |
|
||||||
|
|
||||||
|
## Recommendations
|
||||||
|
|
||||||
|
- Consider adding domain knowledge for new focus areas
|
||||||
@@ -3,24 +3,27 @@ title: Infrastruktur-Übersicht
|
|||||||
category: entities
|
category: entities
|
||||||
tags: [homelab, hardware, overview]
|
tags: [homelab, hardware, overview]
|
||||||
created: "2026-04-28"
|
created: "2026-04-28"
|
||||||
modified: "2026-07-24"
|
modified: "2026-09-26"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Infrastruktur-Übersicht
|
# Infrastruktur-Übersicht
|
||||||
|
|
||||||
## Physikalische Hardware
|
## Physikalische Hardware
|
||||||
|
|
||||||
### Proxmox Cluster (PVE 9.2.3)
|
### Proxmox Cluster (PVE 9.2.20)
|
||||||
- **9 Nodes**, Quorum OK
|
- **8 Nodes**, Quorum OK (proxmox1+2 dauerhaft entfernt)
|
||||||
- Hypervisoren in 10.0.20.x
|
- Hypervisoren in 10.0.20.x
|
||||||
- Siehe [[systems/proxmox-cluster]]
|
- Siehe [[systems/proxmox-cluster]]
|
||||||
|
|
||||||
### Ceph Cluster
|
### Ceph Cluster
|
||||||
- **14 OSDs** (HDD + NVMe/SSD混合)
|
- **17 OSDs** (10 HDD + 7 SSD), all up/in, v20.2.4
|
||||||
|
- HEALTH_WARN (insecure key types — cosmetic)
|
||||||
|
- 4 MONs (px5 leader, px4, a2-1, n5pro), MGR auf n5pro
|
||||||
- Siehe [[systems/ceph-cluster]]
|
- Siehe [[systems/ceph-cluster]]
|
||||||
|
|
||||||
### RKE2 Kubernetes Cluster
|
### RKE2 Kubernetes Cluster
|
||||||
- **6 Nodes** (3 CP + 3 Worker), v1.35.6-rke2r1
|
- **6 Nodes** (3 CP + 3 Worker), v1.35.6-rke2r1
|
||||||
|
- worker-04 currently NotReady (VM 139 offline)
|
||||||
- Alle Nodes schedulable (keine CP Taints)
|
- Alle Nodes schedulable (keine CP Taints)
|
||||||
- Siehe [[systems/rke2-kubernetes]]
|
- Siehe [[systems/rke2-kubernetes]]
|
||||||
|
|
||||||
|
|||||||
@@ -12,6 +12,7 @@
|
|||||||
| `systems/` | Software-Systeme und Services |
|
| `systems/` | Software-Systeme und Services |
|
||||||
| `concepts/` | Abstrakte Patterns & Konventionen |
|
| `concepts/` | Abstrakte Patterns & Konventionen |
|
||||||
| `reference/` | Quick-Lookup Tabellen (IPs, Keys, Ports) |
|
| `reference/` | Quick-Lookup Tabellen (IPs, Keys, Ports) |
|
||||||
|
| `patterns/` | Rekurrente Failure-Mode Patterns + Skill-Impact Tracker |
|
||||||
|
|
||||||
## Entities
|
## Entities
|
||||||
- [[entities/infrastructure]] — Physikalische Hardware, Cluster-Übersicht
|
- [[entities/infrastructure]] — Physikalische Hardware, Cluster-Übersicht
|
||||||
@@ -21,8 +22,8 @@
|
|||||||
- [[entities/health-fitness]] — Gesundheits-Ziele, Ernährung
|
- [[entities/health-fitness]] — Gesundheits-Ziele, Ernährung
|
||||||
|
|
||||||
## Systems
|
## Systems
|
||||||
- [[systems/proxmox-cluster]] — PVE 9.2.3, 9 Nodes, Fluent Bit
|
- [[systems/proxmox-cluster]] — PVE 9.2.20, 8 Nodes, Fluent Bit
|
||||||
- [[systems/ceph-cluster]] — 14 OSDs, EC Pools, bekannte Probleme
|
- [[systems/ceph-cluster]] — 17 OSDs, v20.2.4, 4 MONs, bekannte Probleme
|
||||||
- [[systems/rke2-kubernetes]] — 6 Nodes, ArgoCD, CNPG, Workloads
|
- [[systems/rke2-kubernetes]] — 6 Nodes, ArgoCD, CNPG, Workloads
|
||||||
- [[systems/galera-maxscale]] — 3 Galera + 2 MaxScale, VIP .70
|
- [[systems/galera-maxscale]] — 3 Galera + 2 MaxScale, VIP .70
|
||||||
- [[systems/loki-fluentbit]] — Logging Stack, 40+ Targets
|
- [[systems/loki-fluentbit]] — Logging Stack, 40+ Targets
|
||||||
@@ -30,18 +31,43 @@
|
|||||||
- [[systems/hindsight]] — Semantic Memory, K8s, API
|
- [[systems/hindsight]] — Semantic Memory, K8s, API
|
||||||
- [[systems/monitoring]] — Prometheus, Grafana, HolmesGPT
|
- [[systems/monitoring]] — Prometheus, Grafana, HolmesGPT
|
||||||
- [[systems/seafile]] — cloud.familie-schoen.com, Seafile 13.0.19
|
- [[systems/seafile]] — cloud.familie-schoen.com, Seafile 13.0.19
|
||||||
|
- [[systems/laya]] — 421M Classifier CT152, choice/noul/scale, Session-Pre-Filter
|
||||||
|
- [[systems/noris-ai]] — ai.noris.de LLMs, Embeddings, Image Gen
|
||||||
|
- [[systems/frigate]] — NVR v0.18 CT151, Kameras, MQTT, HA Automation
|
||||||
|
- [[systems/homeassistant]] — HAOS, MQTT, Automations, Notify
|
||||||
|
- [[systems/paperless]] — Dokumentenmgmt, OIDC, OCR via noris AI
|
||||||
|
- [[systems/sarah-hermes]] — VM107, zweites Hermes für Sarah (WebUI :8787, TG geplant)
|
||||||
|
|
||||||
## Concepts
|
## Concepts
|
||||||
- [[concepts/network-architecture]] — 10.0.X.Y Schema, VLANs
|
- [[concepts/network-architecture]] — 10.0.X.Y Schema, VLANs
|
||||||
- [[concepts/gitops-workflow]] — IaC → Git → ArgoCD → Verify
|
- [[concepts/gitops-workflow]] — IaC → Git → ArgoCD → Verify
|
||||||
- [[concepts/credential-policy]] — 1Password, ESO, keine Secrets in Git
|
- [[concepts/credential-policy]] — 1Password, ESO, keine Secrets in Git
|
||||||
- [[concepts/email-organization]] — Rechnungs-Organizer, Himalaya
|
- [[concepts/email-organization]] — Rechnungs-Organizer, Himalaya
|
||||||
|
- [[concepts/memory-layer-architecture]] — 7-Schichten Memory Trennungsregeln
|
||||||
|
|
||||||
## Reference
|
## Reference
|
||||||
- [[reference/ip-map]] — IP → Host → Service Mapping
|
- [[reference/ip-map]] — IP → Host → Service Mapping
|
||||||
- [[reference/ssh-keys]] — Key → Zweck → Fingerprint
|
- [[reference/ssh-keys]] — Key → Zweck → Fingerprint
|
||||||
- [[reference/ports]] — Port → Service → Host
|
- [[reference/ports]] — Port → Service → Host
|
||||||
|
|
||||||
|
## Patterns
|
||||||
|
- [[patterns/_README]] — Konzept und Nutzung des Patterns-Verzeichnisses
|
||||||
|
- [[patterns/galera-ddl-deadlock]] — OPTIMIZE TABLE → TOI Deadlock (PAT-001)
|
||||||
|
- [[patterns/traefik-reload-unreliable]] — reload unreliable, restart statt reload (PAT-002)
|
||||||
|
- [[patterns/1password-rate-limit]] — Account-Level Rate-Limit, ESO stoppen (PAT-003)
|
||||||
|
- [[patterns/vfio-gpu-passthrough]] — softdep + DRM blacklist Race Condition (PAT-004)
|
||||||
|
- [[patterns/finanzblick-sync-waf]] — POST /sync WAF-blocked, UI Modal Sequence (PAT-005)
|
||||||
|
- [[patterns/linkedin-react-input]] — nativeInputSetter, Session Expiry (PAT-006)
|
||||||
|
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
|
||||||
|
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
|
||||||
|
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
|
||||||
|
- [[patterns/pve-oom-frozen-guest]] — Host-OOM friert Gast als Zombie ein, RSS-Kollaps-Signatur (PAT-010)
|
||||||
|
- [[patterns/pve-notification-webhook]] — PVE Webhook→Telegram: Endpoint-Drift, Base64-Trap, Handlebars-Escape (PAT-011)
|
||||||
|
- [[patterns/placement-guardrails]] — 70% RAM-Deckel + Swap-Verbot als Prometheus-Rules (PAT-012)
|
||||||
|
- [[patterns/prometheus-config-rescue]] — grep -vn Gift/Crashloop, LXC-Rootfs-Direct-Mount-Rescue (PAT-013)
|
||||||
|
- [[patterns/ceph-dead-nvme-resurrection]] — PCI-Reset→lvchange→chown-Falle, Weight-Management bei Reintegration (PAT-014)
|
||||||
|
- [[patterns/skill-impact]] — Skill Modification Audit Trail
|
||||||
|
|
||||||
## Memory Layer Architektur
|
## Memory Layer Architektur
|
||||||
| Layer | System | Pfad | Rolle |
|
| Layer | System | Pfad | Rolle |
|
||||||
|-------|--------|------|-------|
|
|-------|--------|------|-------|
|
||||||
|
|||||||
@@ -1,6 +1,174 @@
|
|||||||
# Memory Log
|
# Memory Log
|
||||||
|
|
||||||
## [2026-07-26] fix | Gitea SSH Push Key + Paperless IngressRoute Hostname
|
## [2026-09-30] k8s-cp01-relief-descheduler | Workload-Migration + Descheduler-Deployment
|
||||||
|
- cp-01 bei 93% RAM → hindsight-api + hindsight-postgres + paperless nach worker-05 migriert (cordon/delete/uncordon). cp-01 RAM 93%→43%.
|
||||||
|
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren bis konvergiert.
|
||||||
|
- hindsight-api Live-Pin (Sep-17-Hotfix) war nie in Git → Live-Patch nötig.
|
||||||
|
- Descheduler v0.30.0 deployt (raw Manifests, nicht Helm — Chart hat params/args-Bug + falsche API-Version). Policy: LowNodeUtilization (30/50/20 → 70/70/60), 2m Intervall, nodeFit auf Profile-Ebene.
|
||||||
|
- RBAC-Lektion: Descheduler braucht `list/watch` auf `namespaces` in ClusterRole.
|
||||||
|
- ArgoCD repoURL: `ssh://git@10.0.30.200:22/...` (mit `git@` User-Prefix, wie alle anderen Apps).
|
||||||
|
- Solution Doc + INDEX + Wiki rke2-kubernetes.md aktualisiert. Commits b955d7e→8c1ecff.
|
||||||
|
|
||||||
|
## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur
|
||||||
|
- Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen.
|
||||||
|
- Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate.
|
||||||
|
- SMART-Korrektur: Kingston SFYRDK2000G (osd.2-Träger) ist GESUND (wear 6 %, media_errors 0, spare 100 %, PoH 1469) — frühere "Austausch-Kandidatin"-These RETRAKIERT. Vorfall war Controller-/Fabric-Ebene. Neu beobachten: Temps 70/78 °C, T1-throttle 3×.
|
||||||
|
- PAT-014 erweitert (Smart-Log-Abschnitt), ceph-cluster.md korrigiert, Commit cc2f237.
|
||||||
|
|
||||||
|
## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014)
|
||||||
|
- osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin. *(Korrektur später am selben Tag: SMART-Audit → gesund, kein Austausch; siehe nächster Eintrag.)*
|
||||||
|
- osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend.
|
||||||
|
- ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben.
|
||||||
|
- Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014.
|
||||||
|
|
||||||
|
## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013)
|
||||||
|
- Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün.
|
||||||
|
- Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok.
|
||||||
|
- **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate.
|
||||||
|
- Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit <SHA>.
|
||||||
|
|
||||||
|
## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012)
|
||||||
|
- Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok.
|
||||||
|
- Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom).
|
||||||
|
- Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172.
|
||||||
|
|
||||||
|
## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011)
|
||||||
|
- Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON).
|
||||||
|
- Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200.
|
||||||
|
- Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body.
|
||||||
|
- Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales).
|
||||||
|
|
||||||
|
## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen
|
||||||
|
- proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%.
|
||||||
|
- Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets.
|
||||||
|
- VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP.
|
||||||
|
- Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt.
|
||||||
|
|
||||||
|
## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet
|
||||||
|
- GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md.
|
||||||
|
- Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc.
|
||||||
|
- proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen.
|
||||||
|
- Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert.
|
||||||
|
|
||||||
|
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
|
||||||
|
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
|
||||||
|
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
|
||||||
|
- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX
|
||||||
|
- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared
|
||||||
|
- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr
|
||||||
|
- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return
|
||||||
|
- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010)
|
||||||
|
|
||||||
|
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
|
||||||
|
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
|
||||||
|
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
|
||||||
|
- E2E verifiziert: Login 200, Agent-Antwort via vllm/release/glm-5-2 @ ai.noris.de
|
||||||
|
- PITFALL: hermes-agent pyproject pinnt Core-Deps hinter python_version>='3.14'-Markers → auf Python 3.12 installiert `pip install -e .` NULL Deps ("AIAgent not available"). Fix: manuelle Pin-Installation + Missing-Import-Loop bis `import run_agent`
|
||||||
|
- Backup: täglich 23:00 Job backup-2e8a34e3-66cb (all=1) deckt VM107; manueller Verify TASK OK 48s
|
||||||
|
- Wiki: systems/sarah-hermes.md, ip-map.md erweitert; iac-homelab commits 0d5f017/1ce4816
|
||||||
|
- OFFEN: Telegram-Bot blockiert auf BotFather-Token (Sarah/Dominik)
|
||||||
|
|
||||||
|
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
|
||||||
|
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
|
||||||
|
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
|
||||||
|
- Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min)
|
||||||
|
- Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05
|
||||||
|
- Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2)
|
||||||
|
- Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes)
|
||||||
|
- Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md
|
||||||
|
|
||||||
|
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
|
||||||
|
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
|
||||||
|
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
|
||||||
|
- Gitea Pod + schoenkitchen runner durch Pod-Delete wiederhergestellt
|
||||||
|
- ArgoCD Audit: 17/25 Apps via Gitea SSH, 20/25 Synced+Healthy
|
||||||
|
- Known Issues: `gitea`/`authelia` Health=Progressing (Ingress-LB-IP Gap), `immich` doppelt verwaltet, `gitea-config` Deployment Drift
|
||||||
|
- `DEFAULT_ACTIONS_URL=https://gitea.com` deprecated → Helm override auf `self` empfohlen
|
||||||
|
- Wiki `systems/gitea.md` aktualisiert: Runner Status, Incident, ArgoCD Issues
|
||||||
|
- Solution Doc: `docs/solutions/bug-fixes/2026-09-27-gitea-csi-rbac-volumeattachment-fix.md`
|
||||||
|
|
||||||
|
## [2026-09-27] architecture | Memory Layer Restructuring + Laya Session-Nutzung
|
||||||
|
- USER.md bereinigt: Infra-Fakten entfernt, nur noch User-Preferences/Safety-Rules (1.087/1.375 chars)
|
||||||
|
- MEMORY.md ausgedünnt: 2.145→1.664 chars, alle Infra-Details zeigen auf Wiki-Seiten mit `→Wiki` Pointern
|
||||||
|
- 4 neue Wiki-Seiten: systems/noris-ai, systems/frigate, systems/homeassistant, systems/paperless
|
||||||
|
- Neues Concept: concepts/memory-layer-architecture — Entscheidungsbaum "was wohin gehört"
|
||||||
|
- Hindsight Audit: massiv überladen mit veralteter Infra-Topologie. Going-Forward-Policy: nur noch semantische Pointer + Entscheidungen
|
||||||
|
- LCM gesund: 7.590 messages, 34 DAG nodes, 21.2:1 compression ratio
|
||||||
|
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
|
||||||
|
|
||||||
|
## [2026-09-27] architecture | Laya Email-Organizer Migration + Session-Nutzung
|
||||||
|
- Rechnungen-Organizer Cron `f773f8c23230` von LLM-Agent → `no_agent` Script mit Laya migriert
|
||||||
|
- 12 Kategorien, 3-Schichten-Safety (Confidence-Gate + Subject-Validierung + Move-Erfolg)
|
||||||
|
- Globale Ordner (Rechnungen/Bestellungen/Gutschriften/Gutscheine) statt Monatssortierung
|
||||||
|
- Wiki-Seite `systems/laya.md` erstellt mit Usage Guide für Session-Nutzung
|
||||||
|
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
|
||||||
|
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
|
||||||
|
|
||||||
|
## [2026-09-17] workflow | CT108-Endausbau: Census, Ghost-Router CT99999, Alias-Repair, Stop
|
||||||
|
- **Census (auth):** K8s=25 / CT108=31 / gemeinsam=23 — 22 Tips identisch, dominik/memory=K8s-Superset (enthält CT-Tip de97cd6d), nur-CT=8× PoC-Müll. Archiv: 29 Bare-Bundles 143 MB unter `/home/debian/git-archive/ct108-final/` (2 Failures = legitim leere Repos).
|
||||||
|
- **Ghost-Router enttarnt:** CT99999 (Traefik-LXC auf proxmox7, 10.0.60.10) terminiert TLS für *.familie-schoen.com und forwardet plain HTTP an .203. `git.familie-schoen.com` → .105:3000 WAR der letzte Live-Konsument von CT108; `git.schoen.codes` lief längst per Double-Hop aufs K8s. Fix: gitea-service-Upstream → .203 (Backup `explicit-http.yml.bak-hermes-20260917`) + Alias-Host im K8s-Ingress (Commit `9e9b6ee`). Beide Hostnamen jetzt v1.27.0.
|
||||||
|
- **Stop vollzogen (genehmigt):** `pct stop 108` auf ms-a2-2 (10.0.20.93) — connection-refused-Beweis, Fleet 22/25 Synced, Runner unversehrt. Wiki-Lügen korrigiert (ip-map behauptete „stopped" seit Wochen).
|
||||||
|
- **Lessons:** 1P-SA braucht je Call `--vault`; `op read --reveal` existiert nicht (stdout=Secret); `op item get --reveal` maskiert nur Display (JSON-Captures intakt); `/repos/search` = `{ok,data}`-Envelope; blankes Token erzeugt glaubwürdig LEERE Census (Fast-Fehlentscheidung „K8s hat nur 2 Repos").
|
||||||
|
- Docs: `docs/solutions/workflows/2026-09-17-ct108-full-decommission-census-router-topology.md` · Wiki: systems/gitea.md, reference/ip-map.md
|
||||||
|
- Offen (je Freigabe): Zombie-Secret `argocd-repo-credentials` löschen; immich-Zwillings-App (toter rendered-Dump vom 01.08., Live gehört immich-config) entfernen.
|
||||||
|
|
||||||
|
## [2026-09-17] fix | Merge-Day-Kampagne abgeschlossen: Konvergenz + Autosync-Rennen + Live-State-Arbitrage
|
||||||
|
- **Konvergenz DONE:** Merge `d76d3e8` (fork 46fa168, 24.07.) + Fixups `4987341`/`8c741c5` auf BEIDEN Remotes (CT108 + K8s-Gitea SSH 10.0.30.200). Ahead37/behind63-Narrativ endgültig begraben (Cache-Phantom). Single-Remote-Ziel erreicht: beide Tipps identisch.
|
||||||
|
- **Autosync-Renne verarbeitet (4 Minentypen):** authelia Duplikat-Volume (union-merge) entfernt; homepage `authelia-auth`-Middleware restauriert (Jul-Entscheidung ging nie live); paperless plaintext-OIDC-Secret GELÖSCHT statt Wert-Rollback (ESO-Ownership seit 24.08., Quelle 1P); newborn-App ceph-csi-cephfs eingefroren (autosync-Block aus Git entfernt — Helm-Release 3.17.0 wartet auf Adoption; rbd-chart 3.10.1 weiterhin ungoverned, Backlog).
|
||||||
|
- **3 Sync-Failures via Live-State-Arbitrage geheilt (alle Synced/Healthy @ 8c741c5):** (1) backups: VSC-CRD-Flavor-Falle — RKE2-Addon-CRD ist FLAT (driver/deletionPolicy top-level, `spec:` verboten!), Jul-Files waren für Upstream-Flavor korrekt; beide Files geflattet, cephfs-snapclass erstmals LIVE ERZEUGT. (2) databases: CNPG grow-only — Git auf 100Gi/20Gi hochaligned (Shrink verboten). (3) hindsight: VCT storageClassName immutable (hdd→flash aligniert), Live-Hotfixes gespiegelt (Node-Pin cp-01, ReadinessProbe draußen, secretKeyRef statt Klartext-PW). hindsight-postgres-0 rotierte sauber, PVC 47d Bound, API pollt.
|
||||||
|
- **Push-Learned:** canonical-Push braucht EXPLIZITEN Key (`GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o IdentitiesOnly=yes"`) — kein ~/.ssh/config vorhanden, Default-Key lehnt ab.
|
||||||
|
- **Chronisch (prä-merge, offen):** kube-prometheus-stack Synced/FAILED (CRD-Annotation >256kB); residual OutOfSync gitea-config (Deployment/gitea-runner) + immich-config (SA/CM-Reste) = Alt-Backlog, keine Regression.
|
||||||
|
- **Finale Ordnung (nächste Schritte):** ① ArgoCD-Sources auf ssh://git@10.0.30.200:22 flippen (incl. Secret argocd-repo-credentials) ② CT108 final stoppen (explizite Freigabe Dominik) ③ CSI-Adoption ④ kpstack-CRD-Fix.
|
||||||
|
- Docs: `docs/solutions/workflows/2026-09-17-argocd-autosync-race-four-mine-types.md` + `docs/solutions/bug-fixes/2026-09-17-live-state-arbitration-crd-flavors-hotfix-mirroring.md` · Morgen-Doc (.202-Empfehlung) korrigiert auf .200.
|
||||||
|
|
||||||
|
## [2026-09-17] bug-fix | Ghost-Instance-Regression: CT108-Gitea als ArgoCD-Source + gitea-backup RCA-Quality
|
||||||
|
- gitea-backup-29826930 (03:30Z) failed: BackoffLimitExceeded, 3 Instant-Crashes in 73s. Zwei Auto-RCAs attribuierten auf 1P/ESO-Rate-Limit — STRUKTURELL widerlegt (crashing init-container gitea-files konsumiert keine Secrets; mounted Secrets intakt).
|
||||||
|
- Echter Fund: ArgoCD-Apps (root/proxy/paperless/schoenkitchen/gitea-config) zogen von http://10.0.30.105:3000 = CT108-Ghost (nach Dekommissionierung 24.07. wieder eingeschaltet, stale Mirror ohne Hardening-Commit 2014673) → "Synced" maskierte fehlende failedJobsHistoryLimit-Felder live.
|
||||||
|
- Sofortmassnahme: manueller Rerun gitea-backup-manual-161538 SUCCESS (201,7 MB, S3-Upload verifiziert), Failed-Job gelöscht, Alert clear.
|
||||||
|
- Offen (awaiting owner): ArgoCD-Sources auf ssh://git@10.0.30.202:22 flippen, Repo-Divergenz ahead37/behind63 + tote HTTPS-Auth forensisch, CT108 final stoppen, ESO-Nachtsättigung (00:00Z-Fenster, cf. PAT-003) analysieren.
|
||||||
|
- Doc: `docs/solutions/bug-fixes/2026-09-17-gitops-ghost-instance-regression-masked-hardening.md` · Wiki: systems/gitea (CT108-Status), concepts/gitops-workflow (Single-Remote-Regel)
|
||||||
|
- **Abend-Phase:** SSH-Deny-Rootcause = keine Keys/Tokens im K8s-Gitea-DB registriert (Opfer der 01.08.-Migration) → via 1P-Token (hermes-gitops) re-registriert: User-Key + ArgoCD-Deploy-Key (beide end-to-end verifiziert, push dry-run OK). Echter Branch-Split am 24.07. entdeckt: K8s-Gitea-main=29.07.-Stand (b9d7441, 70 Commits incl. mariadb:11.4-Fix für gitea-backup), CT108/local=38 Commits ab 04.09. Hardening TTL=86400 + Exit-42-Guard via ArgoCD gelanded (a0c9b6b). Heutiger DB-Dump validiert (116 Tables, kompletter Trailer). Nächste Schritte: Historien-Konvergenz → Source-Flip → CT108-Stop (Freigabe).
|
||||||
|
|
||||||
|
## [2026-08-30] retro | Compound Learning Retrospective (Last 30 Days)
|
||||||
|
- Reviewed sessions from Jul 31 – Aug 30, 2026
|
||||||
|
- **4 new solution docs** written by subagents:
|
||||||
|
- `bug-fixes/2026-08-05-ceph-squid-ec-pool-mark-complete-bug.md` — EC pool mark-complete fails in Squid
|
||||||
|
- `bug-fixes/2026-08-06-traefik-cross-namespace-routing-pitfall.md` — Two Traefik 404 incidents
|
||||||
|
- `architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md` — Full RKE2 rebuild sequence
|
||||||
|
- `architecture/2026-08-01-immich-data-loss-no-backup.md` — Ceph redundancy ≠ backup
|
||||||
|
- **2 new patterns** extracted:
|
||||||
|
- PAT-008: Ansible default_ipv6 fact missing on fresh VMs
|
||||||
|
- PAT-009: Stale NBD devices after RBD volume swap
|
||||||
|
- **4 Hindsight entries** indexed with solution summaries
|
||||||
|
- Key themes: K8s cluster disaster recovery, Ceph EC pool unrecoverability, Traefik routing complexity, backup gap identification
|
||||||
|
|
||||||
|
## [2026-08-30] feat | WikiSkill Patterns Directory + Skill-Impact Tracker
|
||||||
|
- Analysed arXiv:2608.27454 (WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution)
|
||||||
|
- **New `patterns/` directory** in LLM Wiki — structured failure-mode patterns inspired by WikiSkill's Wiki Layer
|
||||||
|
- Created 7 initial pattern pages from existing MEMORY.md entries + solution docs:
|
||||||
|
- PAT-001: Galera DDL TOI Deadlock
|
||||||
|
- PAT-002: Traefik reload unreliable
|
||||||
|
- PAT-003: 1Password account-level rate-limit
|
||||||
|
- PAT-004: VFIO GPU passthrough race condition
|
||||||
|
- PAT-005: Finanzblick sync WAF-blocked
|
||||||
|
- PAT-006: LinkedIn React nativeInputSetter
|
||||||
|
- PAT-007: Ceph EC pool + SSD wear-level
|
||||||
|
- **Skill-Impact Tracker** (`patterns/skill-impact.md`) — audit trail for skill modifications (accept/reject history)
|
||||||
|
- **Template** (`patterns/_template.md`) for future pattern creation
|
||||||
|
- **compound-learning skill patched** — added Phase 4.8 (Pattern Extraction) + Phase 4.9 (Skill-Impact Tracking)
|
||||||
|
- Updated index.md with Patterns section
|
||||||
|
- Hindsight: indexed with tags architecture, wikiskill, patterns, skill-evolution
|
||||||
|
|
||||||
|
## [2026-08-13] feat | Memory Architecture Enhancement (arxiv 2602.06052v4)
|
||||||
|
- Implemented 3 new automated memory capabilities from survey paper analysis:
|
||||||
|
- **Decay & Importance Scoring** (`hindsight_decay_scoring.py`): Monthly cron (1st 03:30). Score = log(1+access)/log(51) × 0.5^(age/90d). Archives score<0.15 with 0 accesses. First run: 200 evaluated, 50 archived.
|
||||||
|
- **User Drift Detection** (`user_drift_detection.py`): Weekly cron (Mon 04:00). Compares 14d vs 90d baseline for message length, frustration, corrections, topic shifts, style changes. First run: detected 3 drifts (41% longer msgs, new focus gitops, declining ai_ml).
|
||||||
|
- **Skill Health Check** (`skill_health_check.py`): Weekly cron (Mon 04:30). Scans 187 active SKILL.md files for stubs, broken refs, missing pitfalls, staleness, duplicates. First run: 70 healthy, 117 warnings, 0 critical.
|
||||||
|
- Cron jobs: 112c106c0954, 280458a2b035, 761b6694d99b (all no_agent, deliver=origin)
|
||||||
|
- Wiki updated: systems/hindsight.md (meta-memory automation table expanded)
|
||||||
|
- Hindsight: indexed with tags architecture, memory-system, cron, arxiv-2602.06052
|
||||||
|
|
||||||
|
## [2026-08-13] fix | Seafile Recovery + Paperless Upgrade + Authelia Ingress
|
||||||
- **Gitea SSH access**: No SSH key on Hermes VM was authorized for Gitea. HTTP port 3000 has no external IngressRoute. Generated dedicated ED25519 key (`id_ed25519_gitea-hermes-push`), stored in 1Password (ID: lqaebdtewj5xatvpzfrkag73v4), registered in Gitea as key ID 4 for user dominik. SSH via `10.0.30.202:22` (LoadBalancer) is the reliable push path.
|
- **Gitea SSH access**: No SSH key on Hermes VM was authorized for Gitea. HTTP port 3000 has no external IngressRoute. Generated dedicated ED25519 key (`id_ed25519_gitea-hermes-push`), stored in 1Password (ID: lqaebdtewj5xatvpzfrkag73v4), registered in Gitea as key ID 4 for user dominik. SSH via `10.0.30.202:22` (LoadBalancer) is the reliable push path.
|
||||||
- **Paperless v3 IngressRoute**: Host was `dokumente.familie-schoen.com` instead of `dokumente-neu.familie-schoen.com`. Two-track fix: kubectl patch (immediate) + Git commit `fa4184b` (permanent via ArgoCD self-heal). Backtick escaping in Traefik Host() match requires `--patch-file` not inline `--patch`.
|
- **Paperless v3 IngressRoute**: Host was `dokumente.familie-schoen.com` instead of `dokumente-neu.familie-schoen.com`. Two-track fix: kubectl patch (immediate) + Git commit `fa4184b` (permanent via ArgoCD self-heal). Backtick escaping in Traefik Host() match requires `--patch-file` not inline `--patch`.
|
||||||
- Wiki updated: systems/gitea.md (SSH push key section), reference/ssh-keys.md (new key entry)
|
- Wiki updated: systems/gitea.md (SSH push key section), reference/ssh-keys.md (new key entry)
|
||||||
@@ -27,7 +195,7 @@
|
|||||||
- Velero backups PartiallyFailed for 121 days — all PVCs skipped
|
- Velero backups PartiallyFailed for 121 days — all PVCs skipped
|
||||||
- Fix 1: Created VolumeSnapshotClass `ceph-rbd-snapclass` (rbd.csi.ceph.com)
|
- Fix 1: Created VolumeSnapshotClass `ceph-rbd-snapclass` (rbd.csi.ceph.com)
|
||||||
- Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22)
|
- Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22)
|
||||||
- Gitea SSH user is `git` not `gitea`; git.schoen.codes → Traefik (10.0.30.203), SSH on 10.0.30.202
|
- Gitea SSH user is `git` not `gitea`; git.schoen.codes → Traefik (10.0.30.208), SSH on 10.0.30.202
|
||||||
- New SSH deploy key generated, added to Gitea repo
|
- New SSH deploy key generated, added to Gitea repo
|
||||||
- Fix 3: `features: EnableCSI` must be under `configuration:` in Velero Helm values (not top-level)
|
- Fix 3: `features: EnableCSI` must be under `configuration:` in Velero Helm values (not top-level)
|
||||||
- Must be string, not array — array form breaks Helm template
|
- Must be string, not array — array form breaks Helm template
|
||||||
@@ -122,3 +290,11 @@
|
|||||||
- Created 2 concept pages: gitops-workflow, credential-policy
|
- Created 2 concept pages: gitops-workflow, credential-policy
|
||||||
- Rewrote index.md with new structure + Memory Layer Architecture table
|
- Rewrote index.md with new structure + Memory Layer Architecture table
|
||||||
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
|
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
|
||||||
|
|
||||||
|
## [2026-09-29] update | VM302 Zombie-Recovery nach Migration
|
||||||
|
- VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot
|
||||||
|
- Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause
|
||||||
|
- Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min
|
||||||
|
- Endstand: 3/3 Synced, Primary, MaxScale alle Server Running
|
||||||
|
- Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md
|
||||||
|
- Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)
|
||||||
|
|||||||
@@ -0,0 +1,51 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-003
|
||||||
|
title: "1Password CLI rate-limit is account-level — stop ESO, wait 1 hour"
|
||||||
|
category: tooling
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [rke2-kubernetes]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/architecture/2026-07-24-1password-account-level-rate-limit.md
|
||||||
|
- docs/solutions/bug-fixes/2026-07-14-1password-token-telegram-corruption.md
|
||||||
|
related_skills: [1password-cli]
|
||||||
|
---
|
||||||
|
|
||||||
|
# 1Password CLI rate-limit is account-level — stop ESO, wait 1 hour
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
All 1Password CLI (`op`) calls suddenly fail with rate-limit errors. This affects
|
||||||
|
both Hermes tools and ExternalSecrets Operator (ESO) in K8s simultaneously.
|
||||||
|
The outage lasts approximately 1 hour.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
The 1Password CLI enforces rate limits at the **account level**, not per-token or
|
||||||
|
per-client. When ESO polls aggressively (multiple SecretStores + frequent sync intervals),
|
||||||
|
it exhausts the quota for the ENTIRE account, blocking all other `op` callers including
|
||||||
|
Hermes automation.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
1. **Stop ESO immediately** — scale down the ESO deployment to halt API calls:
|
||||||
|
```bash
|
||||||
|
kubectl scale deploy -n external-secrets external-secrets-operator --replicas=0
|
||||||
|
```
|
||||||
|
2. **Wait 1 hour** — the rate limit resets automatically. No polling needed.
|
||||||
|
3. **Resume ESO** with reduced polling frequency afterward
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Reduce ESO `refreshInterval` to 10+ minutes (not the default 1 minute)
|
||||||
|
- Limit the number of SecretStores that reference 1Password
|
||||||
|
- Consider caching secrets locally to reduce API pressure
|
||||||
|
- Never run `op` in tight loops — always add delays for batch operations
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Hit twice: Jul 2026 (initial discovery) and during token corruption investigation
|
||||||
|
- Solution doc: `docs/solutions/architecture/2026-07-24-1password-account-level-rate-limit.md`
|
||||||
|
- MEMORY.md entry: "1P CLI:3 vaults R/W. Rate-limit=acct-level. Stop ESO, wait 1hr, no polling."
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
# Patterns Directory
|
||||||
|
|
||||||
|
Inspired by [WikiSkill (arXiv:2608.27454)](https://arxiv.org/abs/2608.27454) — this directory
|
||||||
|
contains structured **failure-mode patterns** and **successful strategies** extracted from
|
||||||
|
real operational experience.
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Unlike `systems/` (describes what exists) or `concepts/` (abstract conventions),
|
||||||
|
`patterns/` captures **recurrent problems and their proven solutions** — the "lessons learned"
|
||||||
|
that compound across incidents.
|
||||||
|
|
||||||
|
## Structure
|
||||||
|
|
||||||
|
Each pattern is a standalone Markdown file:
|
||||||
|
|
||||||
|
```
|
||||||
|
patterns/
|
||||||
|
├── _README.md ← this file
|
||||||
|
├── _template.md ← copy this for new patterns
|
||||||
|
├── galera-ddl-deadlock.md
|
||||||
|
├── traefik-reload-unreliable.md
|
||||||
|
├── ...
|
||||||
|
└── skill-impact.md ← audit trail of skill modifications (accept/reject history)
|
||||||
|
```
|
||||||
|
|
||||||
|
## How Patterns Are Born
|
||||||
|
|
||||||
|
1. **Incident occurs** → problem is diagnosed and fixed
|
||||||
|
2. **Compound-learning cycle** runs → solution doc created in `~/docs/solutions/`
|
||||||
|
3. **Pattern extracted** → if the failure mode is recurrent or broadly applicable,
|
||||||
|
a pattern page is created here
|
||||||
|
4. **Wiki Maintainer** (currently manual, future: automated) consolidates evidence
|
||||||
|
from multiple incidents into the pattern page
|
||||||
|
|
||||||
|
## Relationship to Other Layers
|
||||||
|
|
||||||
|
| Layer | Holds | Retrieval |
|
||||||
|
|-------|-------|-----------|
|
||||||
|
| `~/docs/solutions/` | One-time detailed incident docs | `search_files` (keyword) |
|
||||||
|
| `patterns/` (here) | Recurring failure-mode patterns | Browse `index.md`, cross-linked |
|
||||||
|
| MEMORY.md (L0) | Compressed pointer if high-priority | System prompt injection |
|
||||||
|
| Hindsight (L2) | Semantic index of all above | `hindsight_recall` |
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- **One pattern per file** — don't merge unrelated patterns
|
||||||
|
- **Evidence-based** — cite real incidents (link to solution docs or session dates)
|
||||||
|
- **Actionable** — every pattern must have a "Mitigation" or "Prevention" section
|
||||||
|
- **Never delete without redirect** — if a pattern is obsolete, mark `status: superseded`
|
||||||
|
and link to the replacement
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-XXX
|
||||||
|
title: "<short descriptive title>"
|
||||||
|
category: database|infrastructure|storage|networking|tooling|integration|security
|
||||||
|
severity: low|medium|high
|
||||||
|
status: active|superseded
|
||||||
|
first_observed: YYYY-MM
|
||||||
|
last_updated: YYYY-MM-DD
|
||||||
|
related_systems: []
|
||||||
|
related_solution_docs: []
|
||||||
|
related_skills: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# <Title>
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
What goes wrong? What are the observable symptoms?
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
Why does it happen? Trace the actual cause, not just the symptom.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
What was the fix? Include commands/snippets if relevant.
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
How to avoid this in the future? (monitoring, lint rule, convention, etc.)
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
- When was this observed? Which incidents?
|
||||||
|
- Links to solution docs, session IDs, MEMORY entries
|
||||||
@@ -0,0 +1,59 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-008
|
||||||
|
title: "Ansible default_ipv6 fact missing on fresh VMs → lablabs.rke2 role fails"
|
||||||
|
category: tooling
|
||||||
|
severity: medium
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07-21
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [rke2-kubernetes]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/bug-fixes/2026-07-21-ansible-rke2-default-ipv6-bug.md
|
||||||
|
related_skills: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# Ansible default_ipv6 fact missing on fresh VMs → lablabs.rke2 role fails
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
The `lablabs.rke2` Ansible role fails on freshly provisioned VMs with a Jinja2 error:
|
||||||
|
```
|
||||||
|
ansible_facts['default_ipv6']['address']
|
||||||
|
```
|
||||||
|
The `default_ipv6` fact doesn't exist on VMs without IPv6 configured, causing the
|
||||||
|
`meta/argument_specs.yml` validation to crash.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
The `lablabs.rke2` role's `meta/argument_specs.yml` references
|
||||||
|
`ansible_facts['default_ipv6']['address']` with a hardcoded default. On fresh VMs
|
||||||
|
without IPv6, `ansible_facts['default_ipv6']` is undefined → Jinja2 raises
|
||||||
|
`UndefinedError`.
|
||||||
|
|
||||||
|
A `pre_task` setting `default_ipv6` via `combine()` doesn't help because the
|
||||||
|
argument_specs validation runs BEFORE pre_tasks — it re-collects facts and
|
||||||
|
overwrites the pre_task fix.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
Patch `meta/argument_specs.yml` directly in the role:
|
||||||
|
```yaml
|
||||||
|
# Replace:
|
||||||
|
default: "{{ ansible_facts['default_ipv6']['address'] }}"
|
||||||
|
# With:
|
||||||
|
default: "{{ ansible_facts.default_ipv6.address | default(None) }}"
|
||||||
|
```
|
||||||
|
|
||||||
|
Or: enable IPv6 on the target VMs (faster workaround, no role patching needed).
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Pin Ansible roles and patch argument_specs when they assume facts that may not exist
|
||||||
|
- Test roles on fresh VMs without IPv6 before production use
|
||||||
|
- Consider forking the role with the fix upstream
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed during GPU worker provisioning (worker-04/05, Jul 2026)
|
||||||
|
- Solution doc: `docs/solutions/bug-fixes/2026-07-21-ansible-rke2-default-ipv6-bug.md`
|
||||||
|
- Session: @session:default/20260721_115501_78b25032
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# PAT-014 — Dead-NVMe-Resurrection für Ceph-OSD
|
||||||
|
|
||||||
|
## Trigger
|
||||||
|
Ceph-OSD auf externem/non-Corosync-Host startet nicht: KernelDevice errno-5 I/O-Errors beim
|
||||||
|
Label-Lesen, Symlink `/var/lib/ceph/osd/ceph-*/block` ins Leere, VG „verschwindet".
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
NVMe-Controller im PCIe-Fabric gestorben (state=dead, Namespace 0B) — NICHT das Medium selbst.
|
||||||
|
Nach PCI-reset kehrt der Controller oft unter NEUER Nummer zurück (nvme0 → nvme2!), alte
|
||||||
|
Device-Mapper-Tabelle bleibt stale.
|
||||||
|
|
||||||
|
## Resurrection-Sequence (verifizierte Reihenfolge)
|
||||||
|
1. **Korrekte PCI-Addr finden**: über sysfs-Pfad des Namespaces (`/sys/class/block/nvmeXnY/device`),
|
||||||
|
NICHT raten — erster Versuch traf den falschen (gesunden!) Controller.
|
||||||
|
2. `echo 1 > /sys/bus/pci/devices/<ADDR>/remove && echo 1 > /sys/bus/pci/rescan`
|
||||||
|
→ Controller kommt als neue Instanz zurück, Namespaces wieder da.
|
||||||
|
3. `pvscan --cache` → VG/LV wieder sichtbar.
|
||||||
|
4. **Stale DM-Table**: `lvchange -an <vg>/<lv> && lvchange -ay -K <vg>/<lv>`
|
||||||
|
5. **Lesetest**: `dd if=/dev/<vg>/<lv> bs=4M count=8 of=/dev/null` — bei weiteren errno-5:
|
||||||
|
NAND/Media defekt → OSD out lassen, Drive tauschen.
|
||||||
|
6. **Permissions-Falle nach lvchange**: udev setzt Owner root →
|
||||||
|
`chown ceph:ceph /dev/mapper/<dm-name>; chmod 660` (sonst `bdev open: (13) Permission denied`).
|
||||||
|
7. `systemctl reset-failed ceph-osd@<id> && systemctl start ceph-osd@<id>`
|
||||||
|
8. `ceph osd in osd.<id>` + Weight restaurieren.
|
||||||
|
|
||||||
|
## Begleitregeln
|
||||||
|
- Klein/niedergewichtetes OSD niemals mit Weight 1.0 reintegrieren, solange Pools an
|
||||||
|
Full-Ratios kratzen → sonst `backfillfull`/`backfill_toofull`-Blockade (Fall osd.3:
|
||||||
|
re-in@1.0 → 96 % voll instant → 13 PGs blocked; Fix: `crush reweight osd.3 0.05`).
|
||||||
|
- Externe Non-Corosync-Hosts in `/etc/hosts` ALLER PVE-Nodes pflegen (GUI-500
|
||||||
|
„hostname lookup failed"), oder echter DNS-Record.
|
||||||
|
|
||||||
|
## Verified
|
||||||
|
2026-09-30: osd.2 revived (Kingston SFYRDK2000G, PCI 03:00.0, → nvme2), 495 GiB, Weight 1.0;
|
||||||
|
Cluster 17/17 up/in; Degraded 0,78 % fallend. osd.3 stabilized @ weight 0.05.
|
||||||
|
|
||||||
|
## Smart-Log-Abgleich 2026-09-30 (korrigiert frühere Alters-These)
|
||||||
|
Kingston SFYRDK2000G (nvme0n2, trägt osd.2): percentage_used **6 %**, media_errors **0**,
|
||||||
|
available_spare 100 %, unsafe_shutdowns 2, power_on_hours 1469, power_cycles 3.
|
||||||
|
Geschwisterplatte nvme1n1 identisches Profil (6 % wear, 0 media errors).
|
||||||
|
⇒ Drive ist MEDIZINISCH GESUND — der Vorfall war rein Controller-(PCI)-Ebene, kein
|
||||||
|
Media-Verschleiß. **Kein Austausch nötig.** Einziges Beobachtungsfeld: Temperatur 70 °C /
|
||||||
|
Sensor2 78 °C (thermisches Throttling T1 3× aktiviert) — Kühlung prüfen wäre sinnvoll,
|
||||||
|
aber keine Akutgefahr.
|
||||||
@@ -0,0 +1,60 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-007
|
||||||
|
title: "Ceph EC pool unusable + SSD wear-level failing — monitor pg states"
|
||||||
|
category: storage
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [ceph-cluster]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md
|
||||||
|
- docs/solutions/bug-fixes/2026-07-13-ceph-ratio-ordering-constraint.md
|
||||||
|
- docs/solutions/bug-fixes/2026-07-13-pveceph-wal-lvm-device-mapper-limitation.md
|
||||||
|
related_skills: [ceph-cluster-administration, ceph-ec-incomplete-pg-recovery]
|
||||||
|
---
|
||||||
|
|
||||||
|
# Ceph EC pool unusable + SSD wear-level failing — monitor pg states
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
Multiple Ceph failure modes manifest simultaneously:
|
||||||
|
1. EC pool (pool 8/media_ec) becomes unusable — PGs stuck in `incomplete` state
|
||||||
|
2. SSD OSD (osd.3 on px4) reports `Wear Leveling Count` SMART attribute indicating
|
||||||
|
imminent failure
|
||||||
|
3. `ceph health` shows `HEALTH_ERR`
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
These are compounded issues:
|
||||||
|
1. **EC pool incomplete PGs**: After OSD loss, EC pools cannot recover without enough
|
||||||
|
surviving shards. The minimum copies requirement for EC k+m encoding is stricter than
|
||||||
|
replicated pools.
|
||||||
|
2. **SSD wear-level failure**: Consumer-grade SSDs in Ceph OSD duty cycle reach wear limits.
|
||||||
|
Once `Wear Leveling Count` exceeds threshold, the SSD becomes read-only or unreliable.
|
||||||
|
3. **Weight imbalance**: Unequal OSD weights cause CRUSH placement failures, leaving PGs
|
||||||
|
perpetually remapped.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
1. **EC pool recovery**: Use the `ceph-ec-incomplete-pg-recovery` skill — specialized
|
||||||
|
procedure for forcing EC PG recovery after OSD loss
|
||||||
|
2. **Replace failing SSD**: Mark OSD `out` → `destroy` → `zap` → physically replace → recreate
|
||||||
|
3. **Fix weight imbalance**: Equalize `pg_num` → `pgp_num`, set `nopgchange=true`,
|
||||||
|
reweight OSDs proportionally to disk capacity
|
||||||
|
4. **Remove empty CRUSH buckets**: Clean up stale host entries from decommissioned nodes
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Monitor SMART attributes monthly — alert on wear-level > 80%
|
||||||
|
- Never use consumer SSDs for Ceph OSDs in production (use enterprise grade with PLP)
|
||||||
|
- Keep EC pool k+m ratio proportional to available OSDs (don't over-provision redundancy)
|
||||||
|
- Regular `ceph pg dump` audits for stuck/unactive PGs
|
||||||
|
- Separate PBS onto its own tier (CephFS), away from RBD pools
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Multiple incidents Jul 2026 (OSD 0+2 destruction, OSD 3 wear-level, EC pool unusable)
|
||||||
|
- Solution docs: `2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`,
|
||||||
|
`2026-07-13-ceph-ratio-ordering-constraint.md`
|
||||||
|
- MEMORY.md entry: "Ceph:HEALTH_ERR. Pool8/media_ec unusable. px4 SSD(osd.3) WearLevel FAILING."
|
||||||
@@ -0,0 +1,53 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-005
|
||||||
|
title: "Finanzblick sync requires modal sequence — POST /sync is WAF-blocked"
|
||||||
|
category: integration
|
||||||
|
severity: medium
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: []
|
||||||
|
related_solution_docs: []
|
||||||
|
related_skills: [finanzblick-cashflow]
|
||||||
|
---
|
||||||
|
|
||||||
|
# Finanzblick sync requires modal sequence — POST /sync is WAF-blocked
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
Programmatic synchronization with Finanzblick (banking data aggregator) fails when
|
||||||
|
calling the `POST /sync` API endpoint directly. The request is blocked by the WAF
|
||||||
|
(Web Application Firewall), returning 403 or connection reset.
|
||||||
|
|
||||||
|
Historical fetches (without sync) work fine with the `--no-sync` flag, avoiding 2FA.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
Finanzblick's WAF detects and blocks automated POST requests to the sync endpoint
|
||||||
|
that don't originate from the legitimate browser session with proper CSRF tokens
|
||||||
|
and session cookies.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
Sync must be performed via the **UI button + 2FA modal sequence**:
|
||||||
|
1. Navigate to the Finanzblick web interface in a browser
|
||||||
|
2. Click the sync button (UI-triggered, not API)
|
||||||
|
3. Handle the 2FA modal sequence in order:
|
||||||
|
- PIN modal → click OK
|
||||||
|
- AUTH modal → click WEITER
|
||||||
|
- ERR modal → click OK
|
||||||
|
4. Wait for sync completion
|
||||||
|
|
||||||
|
For historical data fetches (no sync needed), use the `--no-sync` flag — this
|
||||||
|
bypasses 2FA entirely.
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Never attempt direct `POST /sync` calls — always use the UI flow
|
||||||
|
- The `finanzblick-cashflow` skill encodes this modal sequence
|
||||||
|
- This skill is USER-OWNED and needs `hermes curator adopt` to manage
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed during Finanzblick cashflow analysis sessions (Jul 2026)
|
||||||
|
- MEMORY.md entry: "FB sync=UI btn+2FA modals(PIN→OK,AUTH→WEITER,ERR→OK). POST /sync=WAF-blocked. --no-sync flag for hist.fetches(no 2FA). fb-cashflow skill=USER-OWNED,needs `hermes curator adopt`."
|
||||||
@@ -0,0 +1,52 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-001
|
||||||
|
title: "OPTIMIZE TABLE on Galera triggers TOI DDL → Deadlock"
|
||||||
|
category: database
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [galera-maxscale]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/bug-fixes/2026-07-23-cnpg-timeline-corruption-rebuild.md
|
||||||
|
related_skills: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# OPTIMIZE TABLE on Galera triggers TOI DDL → Deadlock
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
Running `OPTIMIZE TABLE` on a Galera cluster node causes a cluster-wide stall or deadlock.
|
||||||
|
The table lock propagates via TOI (Total Order Isolation) to all nodes, blocking all writes
|
||||||
|
to ALL tables during the operation.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
Galera executes DDL statements (including `OPTIMIZE TABLE`, `ALTER TABLE`, etc.) in TOI mode.
|
||||||
|
This means the DDL is replicated as a global operation that blocks the entire cluster — not
|
||||||
|
just the target table. For large tables, the rebuild phase can take minutes, causing apparent
|
||||||
|
outages.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
1. **Use RSU before DDL**: Switch the node to Rolling Schema Update (RSU) mode before running
|
||||||
|
`OPTIMIZE TABLE`. This prevents cluster-wide blocking:
|
||||||
|
```
|
||||||
|
SET GLOBAL wsrep_OSU_method = 'RSU';
|
||||||
|
-- run OPTIMIZE TABLE on this node only
|
||||||
|
SET GLOBAL wsrep_OSU_method = 'TOI'; -- restore
|
||||||
|
```
|
||||||
|
2. **Schedule during maintenance window** — even with RSU, the node itself is degraded
|
||||||
|
3. **Consider pt-online-schema-change** for large tables — avoids blocking entirely
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Never run `OPTIMIZE TABLE` or `ALTER TABLE` in production without RSU on Galera
|
||||||
|
- Add this check to DBA runbooks and monitoring alerts
|
||||||
|
- The Galera skill (`mariadb-galera-cluster-administration`) documents this pitfall
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed during Galera cluster administration sessions (Jul 2026)
|
||||||
|
- MEMORY.md entry: "OPTIMIZE TABLE=TOI DDL→Deadlock! RSU davor."
|
||||||
|
- Galera cluster: nodes 300/301/302, VIP .70:3306
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-009
|
||||||
|
title: "Stale NBD devices after RBD volume swap — delete VolumeAttachment + nbd teardown"
|
||||||
|
category: infrastructure
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-08-01
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [rke2-kubernetes, ceph-cluster]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md
|
||||||
|
related_skills: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# Stale NBD devices after RBD volume swap — delete VolumeAttachment + nbd teardown
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
After swapping RBD images for PVCs (e.g. migrating to new Ceph pool or recreating
|
||||||
|
images), pods fail to mount with errors like:
|
||||||
|
```
|
||||||
|
MountVolume.MountDevice failed for volume "pvc-xxx" : rpc error: code = Internal
|
||||||
|
desc = rbd: map failed with error: /dev/nbd0 already in use
|
||||||
|
```
|
||||||
|
|
||||||
|
The NBD device is held by a stale mapping from the old RBD image, even though the
|
||||||
|
new image has the same name.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
When an RBD image is recreated (delete + create with same name), the Ceph CSI
|
||||||
|
driver's NBD mappings from the old image remain active. The Linux NBD layer
|
||||||
|
holds `/dev/nbdX` open, blocking new mounts to the same device path.
|
||||||
|
|
||||||
|
The Kubernetes VolumeAttachment object also references the old volume handle,
|
||||||
|
preventing the CSI driver from cleanly attaching the new volume.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
Three-step teardown procedure:
|
||||||
|
```bash
|
||||||
|
# 1. Delete the VolumeAttachment (allows CSI driver to release)
|
||||||
|
kubectl delete volumeattachment csi-cephfsplugin-<node>-<volume-handle>
|
||||||
|
|
||||||
|
# 2. Disconnect the stale NBD device on the target node
|
||||||
|
ssh <node> 'nbd-client -d /dev/nbd0' # or: qemu-nbd --disconnect /dev/nbd0
|
||||||
|
|
||||||
|
# 3. Restart the CSI node plugin to pick up clean state
|
||||||
|
kubectl delete pod -n kube-system csi-cephfsplugin-<node-id>
|
||||||
|
# (DaemonSet will respawn it)
|
||||||
|
```
|
||||||
|
|
||||||
|
After this, the pod can remount with the new RBD image.
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Before deleting RBD images, ensure all pods using them are scaled to 0
|
||||||
|
- Delete VolumeAttachments BEFORE deleting RBD images
|
||||||
|
- After RBD image recreation, restart CSI plugins on all nodes that had mounts
|
||||||
|
- Document this in the K8s disaster recovery runbook
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed during K8s cluster rebuild (Aug 2026) — 13 RBD volumes needed swapping
|
||||||
|
- Session: @session:default/20260731_113701_74fd2814 (250+ tool calls)
|
||||||
|
- Part of the full cluster rebuild procedure
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-006
|
||||||
|
title: "LinkedIn login via browser — nativeInputSetter required, session expires between nav"
|
||||||
|
category: integration
|
||||||
|
severity: medium
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: []
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/bug-fixes/2026-07-10-ha-token-expiry-browser-session-loss.md
|
||||||
|
related_skills: [linkedin-personal-branding]
|
||||||
|
---
|
||||||
|
|
||||||
|
# LinkedIn login via browser — nativeInputSetter required, session expires between nav
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
Automated LinkedIn login via browser tools fails silently. Form fields appear filled
|
||||||
|
but LinkedIn doesn't recognize the input (login button stays disabled, or submission
|
||||||
|
fails with "invalid credentials").
|
||||||
|
|
||||||
|
Additionally, authenticated sessions expire between page navigations, requiring
|
||||||
|
re-login on almost every navigation step.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
1. **React-controlled inputs**: LinkedIn uses React, which manages input state internally.
|
||||||
|
Simply setting `.value` via DOM manipulation doesn't trigger React's onChange handler.
|
||||||
|
The `nativeInputSetter` approach is required:
|
||||||
|
```javascript
|
||||||
|
const setter = Object.getOwnPropertyDescriptor(
|
||||||
|
window.HTMLInputElement.prototype, 'value'
|
||||||
|
).set;
|
||||||
|
setter.call(inputElement, 'my-value');
|
||||||
|
inputElement.dispatchEvent(new Event('input', { bubbles: true }));
|
||||||
|
```
|
||||||
|
|
||||||
|
2. **Session expiry**: LinkedIn's SPA doesn't persist auth tokens reliably across
|
||||||
|
`browser_navigate` calls (each navigation may reset the JS context).
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
- Always use `browser_console` with `nativeInputSetter` for form fills — NOT `browser_click`
|
||||||
|
followed by typing
|
||||||
|
- Perform login + desired action in a SINGLE browsing session (minimize navigations)
|
||||||
|
- UTF-8 files must be saved without BOM (LinkedIn rejects BOM in posted content)
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- The `linkedin-personal-branding` skill documents this workflow
|
||||||
|
- Never use `browser_click` for LinkedIn form fields
|
||||||
|
- Minimize navigation steps after login
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed during LinkedIn branding sessions (Jul 2026)
|
||||||
|
- MEMORY.md entry: "Login via browser_console nativeInputSetter. NOT browser_click. Session expires between nav. UTF-8 no BOM."
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# PAT-012 — Placement Guardrails (RAM-Deckel + Swap-Verbot)
|
||||||
|
|
||||||
|
**Trigger**: Neue Gastplatzierung oder Kapazitätsfrage auf einem PVE-Node.
|
||||||
|
|
||||||
|
**Policy**:
|
||||||
|
- Committed RAM pro Hypervisor ≤ 70 % physikalisch. Über Deckel = "voll",
|
||||||
|
Guests redistribuieren BEVOR neue Last kommt.
|
||||||
|
- Resident Swap > 100 MiB für >10 min = Politikverstoß → Placements prüfen,
|
||||||
|
nie als Normalzustand akzeptieren.
|
||||||
|
|
||||||
|
**Enforcement (live)**: Gruppe `placement_guardrails` in CT141
|
||||||
|
(`/opt/monitoring/prometheus/rules/alerting_rules.yml`):
|
||||||
|
- `PVEPlacementCeilingBreached` (ratio > 0.70, for 15m, warning)
|
||||||
|
- `HVResidentSwapNonzero` (SwapTotal-Free > 100MiB, for 10m, warning)
|
||||||
|
|
||||||
|
**Vor jeder Platzierung**: Ratio-Query ziehen
|
||||||
|
(`pve_memory_usage_bytes{id=~"node/.*"} / pve_memory_size_bytes{id=~"node/.*"}`)
|
||||||
|
— >70 %-Nodes meiden, Headroom liegt primär auf ms-a2-1/ms-a2-2.
|
||||||
|
|
||||||
|
Baseline-Rollout 30.09.: proxmox3 80,6 / p4 79,3 / p5 75,5 / p6 72,6 %
|
||||||
|
über Deckel (Warnburst erwartet = Rebalancing-Backlog, kein Emergency).
|
||||||
|
|
||||||
|
Ref: docs/solutions/architecture/2026-09-30-placement-guardrails-ram-ceiling.md
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
# PAT-013 — Prometheus Config Rescue (Crashloop + LXC Rootfs Channel)
|
||||||
|
|
||||||
|
**Trigger**: Prometheus crashloop nach Config-Rewrite, oder pct exec unbrauchbar.
|
||||||
|
|
||||||
|
**Iron Rules**:
|
||||||
|
1. NIEMALS `grep -n`-Ausgabe als Rewrite-Source — Nummernpräfixe verseuchen die Datei.
|
||||||
|
Nutze `grep -v` (ohne -n) oder Python/awk.
|
||||||
|
2. YAML validieren BEVOR Container recreate (sonst Crashloop = totale Blindheit).
|
||||||
|
3. Beim Löschen von Targets/IPs ALLE Locations sweeppen — Dubletten leben gern in
|
||||||
|
mehreren Jobs derselben Datei (node_exporter-Liste ≠ blackbox-Liste).
|
||||||
|
|
||||||
|
**Rescue-Kanal (universell für LXC auf RBD)**:
|
||||||
|
```bash
|
||||||
|
# Auf dem Hyper visor, der den CT hostet:
|
||||||
|
lsblk # rbdX finden (size matchen)
|
||||||
|
mount /dev/rbdX /mnt/rescue
|
||||||
|
# Dateien direkt editieren unter /mnt/rescue/opt/...
|
||||||
|
umount /mnt/refuge
|
||||||
|
# im CT: docker compose up -d --force-recreate <svc>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Symptom-Signature Crashloop**: `docker logs prometheus` → "yaml: line N: mapping
|
||||||
|
values are not allowed in this context" = klassisches NN:-Präfix-Gift.
|
||||||
|
|
||||||
|
**Stale-Series-Note**: gelöschte Targets erscheinen sekundenweise weiter in Queries —
|
||||||
|
erst nach Scrape-Zyklus-Gap re-checken, dann Erfolg erklären.
|
||||||
|
|
||||||
|
Ref: docs/solutions/bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
# PAT-011 — PVE Webhook Notifications Pipeline Repair
|
||||||
|
|
||||||
|
**Trigger**: PVE notifications (esp. vzdump) reach Telegram not / test returns 500.
|
||||||
|
|
||||||
|
**Signature diagnosis ladder**:
|
||||||
|
1. `pvesh create /cluster/notifications/targets/<name>/test` — error taxonomy:
|
||||||
|
- `Connection refused` → transport/down (check URL target alive)
|
||||||
|
- `failed to render webhook body` → template broken (syntax/base64)
|
||||||
|
- `http status: 400` → bridge rejected payload (escaping!)
|
||||||
|
2. Listener-Sweep über VLAN: `for ip in .xx…; do /dev/tcp/$ip/port probe; done`
|
||||||
|
3. Byte-Level-Truth via temp mini-sniffer (redirect URL, capture, restore!).
|
||||||
|
|
||||||
|
**Hard rules**:
|
||||||
|
- Mutations an `/etc/pve/notifications.cfg` NUR via `pvesh set` (hand-edits poison
|
||||||
|
global deserialization!). Vorher Snapshot.
|
||||||
|
- `body`/header-values/secrets = **base64 blobs**: `printf '%s' tpl | base64 -w0`.
|
||||||
|
- Handlebars helpers: `{{escape title}}` (Space!), NICHT `{{escape:title}}`.
|
||||||
|
- Immer `{{escape title}}`/`{{escape message}}` verwenden — rohe Interpolation bricht
|
||||||
|
bei Apostrophen/Multiline (Backup-Reports!).
|
||||||
|
- Sniffer-Redirect URL IMMER restaurieren.
|
||||||
|
|
||||||
|
Reference: `docs/solutions/bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md`
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-010
|
||||||
|
title: "PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature)"
|
||||||
|
category: infrastructure
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-09
|
||||||
|
last_updated: 2026-09-30
|
||||||
|
related_systems: [proxmox-cluster, rke2-kubernetes]
|
||||||
|
related_solution_docs: [docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md]
|
||||||
|
related_skills: [rke2-cluster-administration, systematic-debugging]
|
||||||
|
---
|
||||||
|
|
||||||
|
# PAT-010: PVE host OOM-kill freezes guest VM as zombie
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
- K8s node NotReady, kubelet heartbeat stops abruptly
|
||||||
|
- PVE shows VM `running`, but: no ping, no SSH, qemu-guest-agent dead,
|
||||||
|
tap-interface TX counters static
|
||||||
|
- **Signature:** `qm status <vmid> --verbose` → kvm RSS collapses to a
|
||||||
|
tiny fraction (<5%) of assigned RAM — guest kernel no longer touches
|
||||||
|
its memory
|
||||||
|
- Downstream alerts (DS misscheduled, workload CrashLoops) fire, but
|
||||||
|
NOTHING alarms on the actual OOM
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks
|
||||||
|
the largest anon-RSS process = biggest kvm. After TWO consecutive kills of
|
||||||
|
the same guest (HA auto-restarts in between), the revived QEMU comes up
|
||||||
|
but the guest kernel stays inert → zombie VM.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
1. Relieve host pressure FIRST (move movable tenants away) — otherwise
|
||||||
|
the unfreeze re-boots the guest into the same thrash.
|
||||||
|
2. Hard cycle the VM: `qm stop` + `qm start` (respect HA guards).
|
||||||
|
3. Re-query `ha-manager status` afterwards — HA may relocate the VM.
|
||||||
|
4. Uncordon the K8s node; orphan DS pods self-reconcile.
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
- Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/
|
||||||
|
IaC; swap is a buffer, not capacity.
|
||||||
|
- PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently
|
||||||
|
invisible to alerting (noticed only via downstream K8s symptoms).
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
- 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed
|
||||||
|
rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second
|
||||||
|
revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration +
|
||||||
|
hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
# Skill Impact Tracker
|
||||||
|
|
||||||
|
> Inspired by WikiSkill (arXiv:2608.27454) — tracks skill modifications, their validation
|
||||||
|
> outcomes, and acceptance decisions. Prevents repeating failed modifications and provides
|
||||||
|
> an audit trail of skill evolution.
|
||||||
|
|
||||||
|
## How to Use
|
||||||
|
|
||||||
|
When a skill is patched, created, or deleted, append an entry to the table below.
|
||||||
|
The entry records WHAT changed, WHY, and WHETHER it helped (if validated).
|
||||||
|
|
||||||
|
### Entry Format
|
||||||
|
|
||||||
|
```
|
||||||
|
| Date | Skill | Change Type | Description | Validation | Outcome | Ref |
|
||||||
|
```
|
||||||
|
|
||||||
|
- **Change Type**: `created` | `patched` | `deleted` | `adopted`
|
||||||
|
- **Validation**: How was success measured? (`manual` | `tested` | `benchmark` | `n/a`)
|
||||||
|
- **Outcome**: `accepted` | `rejected` | `rolled-back` | `pending`
|
||||||
|
- **Ref**: Session ID or solution doc path
|
||||||
|
|
||||||
|
## Audit Trail
|
||||||
|
|
||||||
|
| Date | Skill | Change Type | Description | Validation | Outcome | Ref |
|
||||||
|
|------|-------|-------------|-------------|------------|---------|-----|
|
||||||
|
| 2026-08-30 | compound-learning | patched | Added WikiMaintainer phase + patterns/ guidance from WikiSkill paper analysis | tested | accepted | this session |
|
||||||
|
| 2026-08-30 | (retrospective) | created | 4 solution docs + 2 new patterns (PAT-008, PAT-009) from 30-day retrospective | tested | accepted | this session |
|
||||||
|
| 2026-08-13 | hindsight | patched | Expanded meta-memory automation table (decay scoring, drift detection, skill health) | tested | accepted | session 2026-08-13 |
|
||||||
|
| 2026-08-13 | (3 cron jobs) | created | Decay scoring, drift detection, skill health check — no_agent cron jobs | tested | accepted | session 2026-08-13 |
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- **Rejected proposals persist here** — they are NOT deleted. Future skill updates can
|
||||||
|
consult this log to avoid repeating failed approaches (the WikiSkill paper showed this
|
||||||
|
is critical for effective evolution).
|
||||||
|
- **Neutral changes** (no improvement, no regression) are recorded as `accepted` with
|
||||||
|
`validation: manual` — they may enable future improvements.
|
||||||
|
- **Rollbacks** are explicitly tracked — if a skill patch caused regressions and was
|
||||||
|
reverted, the entry stays with `outcome: rolled-back`.
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-002
|
||||||
|
title: "Traefik `reload` unreliable after conf.d edits — use `restart`"
|
||||||
|
category: infrastructure
|
||||||
|
severity: medium
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [proxmox-cluster, rke2-kubernetes]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/bug-fixes/2026-07-26-paperless-ingressroute-hostname-fix.md
|
||||||
|
related_skills: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# Traefik `reload` unreliable after conf.d edits — use `restart`
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
After modifying Traefik configuration files (especially dynamic config in `conf.d/`),
|
||||||
|
issuing a `reload` signal (SIGHUP) does not reliably pick up the changes. The old
|
||||||
|
configuration remains active, leading to stale ingress routes, incorrect routing,
|
||||||
|
or 404 errors.
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
Traefik's hot-reload mechanism for file-based dynamic configuration can silently fail to
|
||||||
|
detect changes, especially when:
|
||||||
|
- Files are edited in-place (atomic rename not used)
|
||||||
|
- The file watcher misses events on certain filesystems (e.g., overlayfs, NFS)
|
||||||
|
- rsync is used without `--inplace` (creates temp file + rename, which the watcher may miss)
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
Use `restart` instead of `reload` for Traefik container/service after conf.d edits:
|
||||||
|
```bash
|
||||||
|
# Instead of: docker kill -s HUP traefik (or systemctl reload traefik)
|
||||||
|
# Use:
|
||||||
|
docker compose restart traefik
|
||||||
|
# or: systemctl restart traefik
|
||||||
|
```
|
||||||
|
|
||||||
|
Additionally, when syncing config files via rsync, use `--inplace` to avoid
|
||||||
|
temp-file-rename patterns that confuse file watchers:
|
||||||
|
```bash
|
||||||
|
rsync --inplace -av ./conf.d/ /etc/traefik/conf.d/
|
||||||
|
```
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Always use `restart` (not `reload`) after Traefik config changes in the Traefik CT (99999)
|
||||||
|
- Use `rsync --inplace` when pushing config files to the Traefik host
|
||||||
|
- Document this in deployment runbooks
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed during Paperless v3 IngressRoute hostname fix (Jul 2026)
|
||||||
|
- Traefik runs in CT99999, conf.d directory
|
||||||
|
- MEMORY.md entry: "Traefik CT99999: `reload` unreliable after conf.d edits. Use `restart` or `rsync --inplace`."
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
---
|
||||||
|
pattern_id: PAT-004
|
||||||
|
title: "VFIO GPU passthrough fails — missing softdep + incomplete DRM blacklist"
|
||||||
|
category: infrastructure
|
||||||
|
severity: high
|
||||||
|
status: active
|
||||||
|
first_observed: 2026-07
|
||||||
|
last_updated: 2026-08-30
|
||||||
|
related_systems: [proxmox-cluster, rke2-kubernetes]
|
||||||
|
related_solution_docs:
|
||||||
|
- docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md
|
||||||
|
- docs/solutions/bug-fixes/2026-07-21-amd-gpu-passthrough-rombar.md
|
||||||
|
related_skills: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# VFIO GPU passthrough fails — missing softdep + incomplete DRM blacklist
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
GPU passthrough to a VM fails intermittently or consistently. Symptoms include:
|
||||||
|
- `/dev/dri/renderD128` missing in the guest
|
||||||
|
- `amdgpu` driver not loading in guest
|
||||||
|
- Kernel BUG in host dmesg
|
||||||
|
- GPU device visible in `lspci` but not bound to `vfio-pci`
|
||||||
|
|
||||||
|
## Root Cause
|
||||||
|
|
||||||
|
Two intertwined issues:
|
||||||
|
|
||||||
|
1. **Module loading race condition**: Without `softdep amdgpu pre: vfio-pci`, the amdgpu
|
||||||
|
driver races to bind the GPU before vfio-pci can claim it. Normal bind time ~2.4s;
|
||||||
|
racing bind time ~302s (or never succeeds).
|
||||||
|
|
||||||
|
2. **Incomplete DRM blacklist**: If `drm` and `drm_kms_helper` are not blacklisted, the
|
||||||
|
kernel's DRM subsystem grabs the GPU before vfio-pci, preventing passthrough.
|
||||||
|
|
||||||
|
## Mitigation
|
||||||
|
|
||||||
|
Apply to the PVE host's modprobe config:
|
||||||
|
```bash
|
||||||
|
# /etc/modprobe.d/blacklist-drm.conf
|
||||||
|
blacklist drm
|
||||||
|
blacklist drm_kms_helper
|
||||||
|
|
||||||
|
# /etc/modprobe.d/amdgpu-vfio.conf
|
||||||
|
softdep amdgpu pre: vfio-pci
|
||||||
|
```
|
||||||
|
Then rebuild initramfs and reboot:
|
||||||
|
```bash
|
||||||
|
update-initramfs -u -k all
|
||||||
|
reboot
|
||||||
|
```
|
||||||
|
|
||||||
|
Also ensure `rombar=1` in the VM config for AMD GPUs (PCI ID 1002:xxxx):
|
||||||
|
```
|
||||||
|
# /etc/pve/qemu-server/<VMID>.conf
|
||||||
|
hostpci0: 0000:XX:YY.Z,pcie=1,rombar=1,x-vga=1
|
||||||
|
```
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
- Always configure `softdep` + DRM blacklist BEFORE attempting GPU passthrough on a new node
|
||||||
|
- Add to provisioning scripts (Ansible/Tofu) for any node with AMD GPU intended for passthrough
|
||||||
|
- ROCm 7.x: gfx1150 supported, gfx1036 NOT supported — verify GPU compute ISA before assignment
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
- Observed on ms-a2-1 (race condition, Jul 2026) and worker-05/VM 102 (kernel BUG, rombar)
|
||||||
|
- Solution docs: `2026-07-24-vfio-pci-module-loading-race-condition.md`,
|
||||||
|
`2026-07-21-amd-gpu-passthrough-rombar.md`
|
||||||
|
- MEMORY.md entry: "GPU:2 active—Immich(CT111/n5pro gfx1150)+Frigate(CT120/px7). ROCm7.14=gfx1150 supported,gfx1036 NOT."
|
||||||
+39
-30
@@ -3,71 +3,80 @@ title: IP-Map (Quick Reference)
|
|||||||
category: reference
|
category: reference
|
||||||
tags: [ip, network, reference, quick-lookup]
|
tags: [ip, network, reference, quick-lookup]
|
||||||
created: "2026-07-24"
|
created: "2026-07-24"
|
||||||
modified: "2026-07-24"
|
modified: "2026-09-30"
|
||||||
---
|
---
|
||||||
|
|
||||||
# IP-Map
|
# IP-Map
|
||||||
|
|
||||||
## Proxmox Hosts (10.0.20.x) — 9 Nodes
|
## Proxmox Hosts (10.0.20.x) — 8 Nodes
|
||||||
| IP | Hostname | Node ID | Notes |
|
| IP | Hostname | Node ID | Notes |
|
||||||
|----|----------|---------|-------|
|
|----|----------|---------|-------|
|
||||||
| 10.0.20.20 | proxmox2 | 3 | |
|
|
||||||
| 10.0.20.30 | proxmox3 | 4 | |
|
| 10.0.20.30 | proxmox3 | 4 | |
|
||||||
| 10.0.20.40 | proxmox4 | 2 | |
|
| 10.0.20.40 | proxmox4 | 2 | MON |
|
||||||
| 10.0.20.50 | proxmox5 | 5 | MON, MGR |
|
| 10.0.20.50 | proxmox5 | 5 | MON (leader) |
|
||||||
| 10.0.20.60 | proxmox6 | 6 | |
|
| 10.0.20.60 | proxmox6 | 6 | |
|
||||||
| 10.0.20.70 | proxmox7 | 7 | MON, Traefik CT99999 |
|
| 10.0.20.70 | proxmox7 | 7 | MON, Traefik CT99999 |
|
||||||
| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002 |
|
| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002, MON, MGR (active) |
|
||||||
| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0 |
|
| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0, MON |
|
||||||
| 10.0.20.93 | ms-a2-2 | 1 | OSD 14 (SSD 1.8TB), GPU 1002:13c0 |
|
| 10.0.20.93 | ms-a2-2 | 1 | OSD 14 (SSD 1.8TB), GPU 1002:13c0 |
|
||||||
|
|
||||||
> ⚠️ n5pro (10.0.20.91, nodeid 8) ≠ proxmox7 (10.0.20.70, nodeid 7) — separate Nodes!
|
> ⚠️ n5pro (10.0.20.91, nodeid 8) ≠ proxmox7 (10.0.20.70, nodeid 7) — separate Nodes!
|
||||||
> proxmox1 (10.0.20.10, nodeid 1) wurde dauerhaft entfernt (2026-07-24).
|
> proxmox1 (10.0.20.10) und proxmox2 (10.0.20.20) wurden dauerhaft entfernt.
|
||||||
> Immer `pvecm nodes` für kanonische Liste prüfen.
|
> Immer `pvecm nodes` für kanonische Liste prüfen.
|
||||||
|
|
||||||
## Kubernetes Nodes (10.0.30.5x-6x)
|
## Kubernetes Nodes (10.0.30.5x-6x)
|
||||||
| IP | Node | Role |
|
| IP | Node | Role | Status |
|
||||||
|----|------|------|
|
|----|------|------|--------|
|
||||||
| 10.0.30.51 | cp-01 | Control Plane |
|
| 10.0.30.51 | cp-01 | Control Plane | Ready |
|
||||||
| 10.0.30.52 | cp-02 | Control Plane |
|
| 10.0.30.52 | cp-02 | Control Plane | Ready |
|
||||||
| 10.0.30.53 | cp-03 | Control Plane |
|
| 10.0.30.53 | cp-03 | Control Plane | Ready |
|
||||||
| 10.0.30.63 | worker-01 | Worker |
|
| 10.0.30.61 | worker-01 | Worker | Ready |
|
||||||
| 10.0.30.64 | worker-04 | Worker (GPU ✅) |
|
| 10.0.30.64 | worker-04 | Worker (GPU) | **NotReady** |
|
||||||
| 10.0.30.65 | worker-05 | Worker (GPU defekt) |
|
| 10.0.30.65 | worker-05 | Worker (GPU) | Ready |
|
||||||
|
|
||||||
## Database Layer (10.0.30.7x-8x)
|
## Database Layer (10.0.30.7x-8x)
|
||||||
| IP | Host | Service |
|
| IP | Host | Service |
|
||||||
|----|------|---------|
|
|----|------|---------|
|
||||||
| 10.0.30.70 | — | MaxScale VIP (keepalived) :3306 |
|
| 10.0.30.70 | — | MaxScale VIP (keepalived) :3306 |
|
||||||
| 10.0.30.71 | VM300 | Galera db1 (ms-a2-1) |
|
| 10.0.30.71 | VM300 | Galera db1 (n5pro) |
|
||||||
| 10.0.30.72 | VM301 | Galera db2 (proxmox3) |
|
| 10.0.30.72 | VM301 | Galera db2 (proxmox6) |
|
||||||
| 10.0.30.73 | VM302 | Galera db3 (proxmox6) |
|
| 10.0.30.73 | VM302 | Galera db3 (ms-a2-2) |
|
||||||
| 10.0.30.81 | VM310 | MaxScale-01 Admin :8989 |
|
| 10.0.30.81 | VM310 | MaxScale-01 Admin :8989 |
|
||||||
| 10.0.30.82 | VM311 | MaxScale-02 Standby |
|
| 10.0.30.82 | VM311 | MaxScale-02 Standby |
|
||||||
|
|
||||||
## Infrastructure VMs (10.0.30.x)
|
## Infrastructure VMs (10.0.30.x)
|
||||||
| IP | Host | Service |
|
| IP | Host | Service |
|
||||||
|----|------|---------|
|
|----|------|---------|
|
||||||
|
| 10.0.30.66 | sarah-hermes (VM107, n5pro) | Sarahs Hermes: WebUI :8787 (LAN-only, pw) + TG-Bot (geplant) |
|
||||||
| 10.0.30.99 | CT111 | Immich (n5pro), Migration zu K8s geplant |
|
| 10.0.30.99 | CT111 | Immich (n5pro), Migration zu K8s geplant |
|
||||||
| 10.0.30.100 | CT134 | InfluxDB (Migration zu K8s) |
|
| 10.0.30.100 | ubuntu | Physischer Ubuntu-Node (Ceph OSDs, ZFS pool01_n2_redundant) |
|
||||||
| 10.0.30.105 | — | ~~Gitea~~ DECOMMISSIONED (CT108, stopped) |
|
| 10.0.30.105 | — | ~~Gitea~~ CT108 GESTOPPT 2026-09-17 (Archive: /home/debian/git-archive/ct108-final/) |
|
||||||
| 10.0.30.124 | VM200 | IaC Runner (DHCP), SSH `debian` |
|
| 10.0.30.124 | VM200 | IaC Runner (DHCP), SSH `debian` |
|
||||||
| 10.0.30.141 | CT141 | Monitoring (Prometheus/Grafana) |
|
| 10.0.30.141 | CT141 | Monitoring (Prometheus/Grafana) |
|
||||||
| 10.0.30.145 | CT145 | Voice Pipeline (PJSIP :5060) |
|
| 10.0.30.145 | CT145 | Voice Pipeline (PJSIP :5060) |
|
||||||
|
|
||||||
## K8s LoadBalancers (10.0.30.2xx)
|
## K8s LoadBalancers (10.0.30.2xx) — Cilium LB IPAM, all pinned
|
||||||
| IP | Service | Namespace |
|
| IP | Service | Namespace | Port |
|
||||||
|----|---------|-----------|
|
|----|---------|-----------|------|
|
||||||
| 10.0.30.201 | Hindsight API :9177 | hindsight |
|
| 10.0.30.200 | Gitea SSH | gitea | 22 |
|
||||||
| 10.0.30.204 | InfluxDB :8086 | influxdb |
|
| 10.0.30.201 | Grafana | monitoring | 3000 |
|
||||||
| 10.0.30.205 | Grafana | monitoring |
|
| 10.0.30.202 | Prometheus | monitoring | 9090 |
|
||||||
| 10.0.30.206 | Prometheus :9090 | monitoring |
|
| 10.0.30.203 | Traefik (K8s Ingress) | kube-system | 80/443 |
|
||||||
| 10.0.30.207 | Loki :3100 | logging |
|
| 10.0.30.204 | InfluxDB | influxdb | 8086 |
|
||||||
|
| 10.0.30.205 | PostgreSQL (CNPG rw LB) | postgres | 5432 |
|
||||||
|
| 10.0.30.208 | Hindsight API | hindsight | 9177 |
|
||||||
|
|
||||||
|
## External CTs (10.0.30.xxx)
|
||||||
|
| IP | CT/VM | Service | Notes |
|
||||||
|
|----|-------|---------|-------|
|
||||||
|
| 10.0.30.99 | CT111 | Seafile v13.0.19 (Docker) | cloud.familie-schoen.com, routed via K8s Traefik IngressRoute |
|
||||||
|
|
||||||
|
> All IPs pinned via `io.cilium/lb-ipam-ips` annotation. Pool: 10.0.30.200–250.
|
||||||
|
|
||||||
## DMZ / Reverse Proxies (10.0.60.x)
|
## DMZ / Reverse Proxies (10.0.60.x)
|
||||||
| IP | Host | Service |
|
| IP | Host | Service |
|
||||||
|----|------|---------|
|
|----|------|---------|
|
||||||
| 10.0.60.10 | CT9999 | Traefik Reverse Proxy (root/[REDACTED]) |
|
| 10.0.60.10 | CT99999 | Traefik Outer-Proxy (root/[REDACTED]); Terminiert *.familie-schoen.com + leitet schoen.codes als Plain-HTTP an 10.0.30.203 |
|
||||||
|
|
||||||
## Related
|
## Related
|
||||||
- [[concepts/network-architecture]]
|
- [[concepts/network-architecture]]
|
||||||
|
|||||||
+1
-1
@@ -23,7 +23,7 @@ modified: "2026-07-24"
|
|||||||
| 9090 | Prometheus | 10.0.30.206 | LB |
|
| 9090 | Prometheus | 10.0.30.206 | LB |
|
||||||
| 9093 | Alertmanager | K8s internal | |
|
| 9093 | Alertmanager | K8s internal | |
|
||||||
| 9095 | HolmesGPT Adapter | K8s internal | Alertmgr → HolmesGPT |
|
| 9095 | HolmesGPT Adapter | K8s internal | Alertmgr → HolmesGPT |
|
||||||
| 9177 | Hindsight API | 10.0.30.201 | LB, NOT localhost |
|
| 9177 | Hindsight API | 10.0.30.208 | LB, NOT localhost |
|
||||||
| 9221 | PVE Exporter | 10.0.30.141 | |
|
| 9221 | PVE Exporter | 10.0.30.141 | |
|
||||||
| 9283 | Ceph Prometheus | 10.0.20.91 | active mgr |
|
| 9283 | Ceph Prometheus | 10.0.20.91 | active mgr |
|
||||||
| 9345 | RKE2 Server URL | 10.0.30.50 | cp-01 |
|
| 9345 | RKE2 Server URL | 10.0.30.50 | cp-01 |
|
||||||
|
|||||||
+47
-52
@@ -3,37 +3,44 @@ title: Ceph Cluster
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [ceph, storage, rbd, ec-pool, osd]
|
tags: [ceph, storage, rbd, ec-pool, osd]
|
||||||
created: "2026-07-24"
|
created: "2026-07-24"
|
||||||
modified: "2026-07-25"
|
modified: "2026-09-26"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Ceph Cluster
|
# Ceph Cluster
|
||||||
|
|
||||||
## Overview
|
## Overview
|
||||||
- **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
|
- **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
|
||||||
- **Health**: HEALTH_OK (recovery complete after OSD 0+2 drain, 0.3% misplaced settling)
|
- **Health**: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed)
|
||||||
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH 2026-07-25), 3 MONs (proxmox5/7/4), MGR on proxmox5
|
- **Version**: 20.2.4 (tentacle) — all 17 OSDs
|
||||||
- **OSDs**: 13 (8 SSD, 5 HDD), all up/in — OSDs 0+2 destroyed+purged 2026-07-25
|
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1)
|
||||||
- **Capacity**: ~22 TiB total, 6.0 TiB used
|
- **OSDs**: 17 (10 HDD, 7 SSD), all up/in
|
||||||
|
- **Capacity**: ~33 TiB total, 9.0 TiB used, 24 TiB avail
|
||||||
|
- **Pools**: 13 pools, 533 PGs (532 active+clean, 1 scrubbing)
|
||||||
|
|
||||||
## OSD Layout
|
## OSD Layout
|
||||||
|
|
||||||
| OSD | Class | Size | Host | Reweight | Notes |
|
| OSD | Class | Size | Host | Reweight | Notes |
|
||||||
|-----|-------|------|------|----------|-------|
|
|-----|-------|------|------|----------|-------|
|
||||||
| 0 | ssd | 188 GB | proxmox2 | — | **DESTROYED 2026-07-25** (92% wear) |
|
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
|
||||||
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | Large HDD |
|
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
|
||||||
| 2 | ssd | 233 GB | proxmox2 | — | **DESTROYED 2026-07-25** (slow ops, 81% full) |
|
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
|
||||||
| 3 | ssd | 238 GB | proxmox4 | 0.95 | |
|
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
|
||||||
| 4 | ssd | 233 GB | proxmox3 | 0.95 | |
|
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
|
||||||
| 5 | ssd | 238 GB | proxmox5 | 0.90 | 80% full |
|
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
|
||||||
| 6 | hdd | 2.8 TiB | ubuntu | 1.0 | Large HDD |
|
| 7 | hdd | 931 GB | proxmox7 | 0.95 | |
|
||||||
| 7 | hdd | 500 GB | proxmox7 | 1.0 | Was 0.80, reweighted 2026-07-24 |
|
| 8 | hdd | 3.6 TiB | ubuntu | 1.0 | |
|
||||||
| 8 | hdd | 2.8 TiB | ubuntu | 1.0 | BlueFS spillover |
|
|
||||||
| 9 | ssd | 1.9 TiB | n5pro | 1.0 | |
|
| 9 | ssd | 1.9 TiB | n5pro | 1.0 | |
|
||||||
| 10 | hdd | 300 GB | proxmox6 | 1.0 | Very small HDD |
|
| 10 | hdd | 931 GB | proxmox6 | 0.95 | |
|
||||||
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | Large HDD |
|
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | |
|
||||||
| 12 | ssd | 1.9 TiB | n5pro | 1.0 | |
|
| 12 | ssd | 1.9 TiB | n5pro | 1.0 | |
|
||||||
| 13 | ssd | 1.8 TiB | ms-a2-1 | 1.0 | |
|
| 13 | ssd | 1.8 TiB | ms-a2-1 | 0.95 | |
|
||||||
| 14 | ssd | 1.8 TiB | ms-a2-2 | 1.0 | New 2026-07-24, nvme0n1 |
|
| 14 | ssd | 1.8 TiB | ms-a2-2 | 0.95 | |
|
||||||
|
| 15 | ssd | 1.8 TiB | ubuntu | 1.0 | New |
|
||||||
|
| 17 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
|
||||||
|
| 18 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
|
||||||
|
|
||||||
|
> OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD.
|
||||||
|
> OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded).
|
||||||
|
|
||||||
## Pools
|
## Pools
|
||||||
|
|
||||||
@@ -44,7 +51,7 @@ modified: "2026-07-25"
|
|||||||
| 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on |
|
| 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on |
|
||||||
| 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | |
|
| 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | |
|
||||||
| 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on |
|
| 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on |
|
||||||
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true (was 120, equalized to 112) |
|
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true |
|
||||||
| 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB |
|
| 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB |
|
||||||
| 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites |
|
| 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites |
|
||||||
| 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | |
|
| 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | |
|
||||||
@@ -58,46 +65,34 @@ modified: "2026-07-25"
|
|||||||
|
|
||||||
## Known Issues
|
## Known Issues
|
||||||
|
|
||||||
### Weight Imbalance Causing Placement Failures (2026-07-24)
|
### HEALTH_WARN: Insecure Key Types (2026-09-26)
|
||||||
HDD hosts have extreme weight disparity: n5pro=10.15TB, ubuntu=5.49TB, proxmox7=0.50TB, proxmox6=0.30TB.
|
Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4.
|
||||||
CRUSH host-level selection (rule 1) often picks only 2 of 4 HDD hosts → up sets with 2 OSDs instead of 3.
|
Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not already set).
|
||||||
Result: PGs stuck in `active+clean+remapped` because up set < min_size.
|
|
||||||
|
|
||||||
**Mitigation (2026-07-24)**:
|
### Small SSDs causing reweightdown
|
||||||
1. Reweighted osd.7 from 0.80 → 1.0 → fixed EC pool 8.3d (NONE → osd.7)
|
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
|
||||||
2. Equalized pool 6 pg_num 120 → 112 + nopgchange=true
|
Consider removing from CRUSH or replacing with larger drives.
|
||||||
3. Manual pg-upmap for stuck PGs: 5.13 → [1,6,7], 6.6c → [11,8,7], 6.58 → [11,6,7]
|
|
||||||
4. All `clean+remapped` eliminated. Triggered rebalancing wave (43 PGs backfilling at 26 MiB/s).
|
|
||||||
|
|
||||||
**Long-term**: Small HDDs (osd.7 0.5TB, osd.10 0.3TB) cause CRUSH placement failures. Replace with larger disks or create separate CRUSH root for large HDDs only.
|
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
|
||||||
|
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
|
||||||
|
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
|
||||||
|
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
|
||||||
|
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
|
||||||
|
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
|
||||||
|
thermal throttle T1 3×.**
|
||||||
|
|
||||||
### Pool 6 pg_num/pgp_num Mismatch (Fixed 2026-07-24)
|
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
|
||||||
Pool hdd_disk had pg_num=120, pgp_num=112 (autoscaler reducing to 32).
|
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
|
||||||
Equalized pg_num to 112. Set nopgchange=true to prevent further autoscaler interference.
|
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
|
||||||
|
|
||||||
### BlueFS Spillover on osd.8
|
### worker-04 (VM 139) NotReady in K8s
|
||||||
osd.8 spilled 128KiB metadata from db device (2.1GiB of 30GiB) to slow device.
|
Node offline — not a Ceph issue but affects Ceph CSI attachments.
|
||||||
Cosmetic warning, no data risk. Fix: `ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-8`
|
|
||||||
|
|
||||||
### Slow Operations on osd.2 and osd.7
|
|
||||||
osd.2 (81% full, fragmentation 0.80) and osd.7 (small HDD) experience slow BlueStore ops.
|
|
||||||
osd.2 NVMe has 92% wear — candidate for replacement.
|
|
||||||
|
|
||||||
### osd.0 NVMe Wear
|
|
||||||
92% Wear, Critical Warning → Austausch planen.
|
|
||||||
|
|
||||||
### EC Pool k=4+m=1 — No Rebalance Headroom
|
|
||||||
With 5 OSDs kein Rebalance Headroom. Siehe Solution Doc: `docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`
|
|
||||||
|
|
||||||
## RBD Management
|
|
||||||
- Proxmox RBD Double-Mount Deadlock Pitfall: Niemals `pct mount` und `pct exec` gleichzeitig auf demselben Container
|
|
||||||
- Siehe Solution Doc: `docs/solutions/bug-fixes/2026-07-23-proxmox-rbd-double-mount-deadlock.md`
|
|
||||||
|
|
||||||
## Access
|
## Access
|
||||||
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.92`
|
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50`
|
||||||
- Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs`
|
- Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs`
|
||||||
- Mon nodes: proxmox5, proxmox7, proxmox4
|
- Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro
|
||||||
- Mgr: proxmox5 (active)
|
- Mgr: n5pro (active)
|
||||||
|
|
||||||
## Related Skills
|
## Related Skills
|
||||||
- `ceph-cluster-administration` (devops)
|
- `ceph-cluster-administration` (devops)
|
||||||
|
|||||||
@@ -0,0 +1,62 @@
|
|||||||
|
---
|
||||||
|
title: "Frigate NVR"
|
||||||
|
category: systems
|
||||||
|
tags: [frigate, nvr, camera, ai, mqtt, proxmox]
|
||||||
|
created: "2026-09-27"
|
||||||
|
modified: "2026-09-27"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Frigate NVR
|
||||||
|
|
||||||
|
> Frigate v0.18.0 auf Proxmox LXC CT151. KI-gestützte Objekterkennung für Kameras Einfahrt + Terrasse.
|
||||||
|
|
||||||
|
## Infrastruktur
|
||||||
|
- **Container:** CT151 auf proxmox6
|
||||||
|
- **IP:** `10.0.30.104:5000`
|
||||||
|
- **Version:** v0.18.0-
|
||||||
|
- **Detector:** OpenVINO CPU (3.2 FPS)
|
||||||
|
- **go2rtc:** v1.9.14
|
||||||
|
- **GPU:** Intel iGPU `/dev/dri/renderD128` (gid=993)
|
||||||
|
|
||||||
|
## Kameras
|
||||||
|
| Name | IP | Substream | Detect FPS |
|
||||||
|
|------|----|-----------|------------|
|
||||||
|
| Einfahrt | `10.0.50.102` | Unterstream | 5.1 fps, 1.0 det |
|
||||||
|
| Terrasse | `10.0.50.103` | Unterstream | 5.0 fps, 2.2 det |
|
||||||
|
|
||||||
|
## MQTT
|
||||||
|
- **Broker:** `10.0.30.10:1883` (Mosquitto auf HA)
|
||||||
|
- **Prefix:** `frigate`
|
||||||
|
- **User:** `frigate`
|
||||||
|
- **Topic:** `frigate/events`
|
||||||
|
|
||||||
|
## Frigate Plus Model
|
||||||
|
- `plus://709bb8ad097786a7f2a37e27684c79a9`
|
||||||
|
- Benötigt `PLUS_API_KEY` (Env-Var in `/config/frigate.env`)
|
||||||
|
|
||||||
|
## Features
|
||||||
|
- Face Recognition (groß): Sarah, Dominik, DHL1, Amazon1, Cleo, Lisa, Annemarie, Marcel Schnitzer, Uli, Eva, Postbote, Fr. Schnitzer
|
||||||
|
- License Plate Recognition (CPU, threshold 0.5)
|
||||||
|
- Semantic Search (groß)
|
||||||
|
- Record: alerts+detections, retain 30 days
|
||||||
|
|
||||||
|
## HA Automation
|
||||||
|
- **ID:** `1714065780832` in `/config/automations.yaml`
|
||||||
|
- **Mode:** `single`, cooldown 600s
|
||||||
|
- **Dedup:** `input_text.frigate_last_event_id` speichert letzten Event
|
||||||
|
- **Flow:** MQTT trigger → 10s delay → snapshot → notify → LLM Vision (optional)
|
||||||
|
- **Labels:** person, car, license_plate, face, dog, cat, amazon, ups, package
|
||||||
|
- **Sublabel-Extraktion:** `sub[0]` (Jinja, Array-Leck-Fix)
|
||||||
|
|
||||||
|
## Bekannte Issues
|
||||||
|
- Stationäre Autos triggerten wiederholt "new" Events → Flood → gefixt mit cooldown+dedup
|
||||||
|
- MQTT ACL blockierte HA-Subscription (User `opendtu` hatte keine Subscribe-Rechte) → gefixt
|
||||||
|
- v0.18 Breaking Changes: `clean_copy` entfernt, `license_plate.mask` Format geändert (list→dict)
|
||||||
|
|
||||||
|
## Old Container
|
||||||
|
- CT120 (frigate): gestoppt, wartet auf Deletion
|
||||||
|
|
||||||
|
## Related
|
||||||
|
- [[systems/homeassistant]] — MQTT, Automations, Notifications
|
||||||
|
- [[systems/noris-ai]] — LLM Vision für Event-Klassifizierung
|
||||||
|
- [[entities/infrastructure]] — CT151 auf proxmox6
|
||||||
@@ -35,6 +35,11 @@ modified: "2026-07-24"
|
|||||||
- VM300/301/302 (Galera): Fluent Bit aktiv
|
- VM300/301/302 (Galera): Fluent Bit aktiv
|
||||||
- VM310/311 (MaxScale): Fluent Bit aktiv
|
- VM310/311 (MaxScale): Fluent Bit aktiv
|
||||||
|
|
||||||
|
## Bekannte Issues
|
||||||
|
- **2026-09-29:** VM302 Zombie-State nach Migration (QEMU running, Services tot, QGA down). Behoben via Stop+Start; SST-Rejoin ~2-3min. Siehe `docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md`
|
||||||
|
- **QEMU-Cmdline-Falle:** `-incoming unix:/run/qemu-server/NNN.migrate -S` bleibt nach Migration im Prozess-String — KEIN Beweis für "paused". Immer `qm monitor <vmid> <<< 'info status'` prüfen.
|
||||||
|
- **HA-Guard:** `ha-manager disable` existiert nicht. Bei Resource-State `request_stop` geht `qm stop` trotzdem durch.
|
||||||
|
|
||||||
## Related Skills
|
## Related Skills
|
||||||
- `mariadb-galera-cluster-administration` (devops)
|
- `mariadb-galera-cluster-administration` (devops)
|
||||||
|
|
||||||
|
|||||||
+22
-6
@@ -3,18 +3,18 @@ title: Gitea (Git Server + CI)
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [gitea, git, ci, actions]
|
tags: [gitea, git, ci, actions]
|
||||||
created: "2026-07-24"
|
created: "2026-07-24"
|
||||||
modified: "2026-07-26"
|
modified: "2026-09-27"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Gitea (Git Server + CI)
|
# Gitea (Git Server + CI)
|
||||||
|
|
||||||
## Instanz
|
## Instanz
|
||||||
- **URL:** git.schoen.codes (K8s Ingress, Traefik → gitea-http ClusterIP)
|
- **URL:** git.schoen.codes (K8s Ingress, Traefik → gitea-http ClusterIP)
|
||||||
- **Internal:** gitea-internal.gitea:3000 (K8s ClusterIP 10.43.5.203)
|
- **Internal:** gitea-http ClusterIP (headless, `None`) — Port 3000/TCP, via Ingress/Traefik erreichbar
|
||||||
- **SSH:** 10.0.30.202:22 (Cilium LoadBalancer, gitea-ssh service)
|
- **SSH:** 10.0.30.200:22 (Cilium LoadBalancer, gitea-ssh service, NodePort 31441) — NICHT .202 (alter Wiki-Fehler)
|
||||||
- **Version:** 1.27.0 (K8s Helm chart v12.7.0)
|
- **Version:** 1.27.0 (K8s Helm chart v12.7.0)
|
||||||
- **Privat:** NIEMALS öffentlich machen (enthält Secrets)
|
- **Privat:** NIEMALS öffentlich machen (enthält Secrets)
|
||||||
- **CT108 (alte Gitea, 10.0.30.105): DECOMMISSIONED** — gestoppt am 2026-07-24
|
- **CT108 (alte Gitea, 10.0.30.105, v1.25.5): GESTOPPT am 2026-09-17 (genehmigt).** War bis zum Abend NICHT gestoppt trotz älterer Wiki-Claims — Live-Checks (v1.25.5-API-Antwort) schlugen die Papierlage. Vor dem Stop: Full-Census (31 Repos auth, 23 gemeinsam — 22 Tips identisch, dominik/memory K8s-Superset), 29 Bare-Bundle-Archive (143 MB) unter `/home/debian/git-archive/ct108-final/`. Letzter Live-Konsument war die Route `git.familie-schoen.com` im CT99999-Traefik (Upstream .105:3000) — umgebogen auf .203 (K8s), Alias-Host im K8s-Ingress ergänzt (Commit 9e9b6ee). Seither servieren BEIDE Hostnamen v1.27.0.
|
||||||
|
|
||||||
## Repositories
|
## Repositories
|
||||||
| Repo | Zweck | Clone |
|
| Repo | Zweck | Clone |
|
||||||
@@ -28,12 +28,15 @@ modified: "2026-07-26"
|
|||||||
- Ed25519 SSH-Key als Gitea Deploy Key (read-only)
|
- Ed25519 SSH-Key als Gitea Deploy Key (read-only)
|
||||||
- ArgoCD Secret `argocd-repo-k8s-gitea` mit SSH Private Key
|
- ArgoCD Secret `argocd-repo-k8s-gitea` mit SSH Private Key
|
||||||
- `known_hosts` ConfigMap mit Gitea SSH Hostkey
|
- `known_hosts` ConfigMap mit Gitea SSH Hostkey
|
||||||
- 15 ArgoCD Apps via SSH (`ssh://gitea@git.schoen.codes:22/dominik/iac-homelab.git`)
|
- 17 ArgoCD Apps via SSH (`ssh://git@10.0.30.200:22/dominik/iac-homelab.git`)
|
||||||
|
- 25 ArgoCD Apps gesamt (Stand Sep 2026)
|
||||||
|
|
||||||
## Gitea Actions CI
|
## Gitea Actions CI
|
||||||
- **Primary Runner:** vm200-host (ID 73, act_runner v0.2.13, host backend) — WORKING
|
- **Primary Runner:** vm200-host (ID 73, act_runner v0.2.13, host backend) — WORKING
|
||||||
- **K8s Runner:** gitea-runner deployment in gitea namespace (DinD sidecar + act_runner v0.2.13) — registered but NO TASK ASSIGNMENT (Gitea 1.27.0 bug)
|
- **K8s Runner (gitea NS):** gitea-runner deployment (DinD sidecar + act_runner v0.2.13) — RUNNING, declare erfolgreich (Sep 2026). Vorher 4 Tage Init-crash wegen CSI RBAC Issue (siehe unten).
|
||||||
|
- **K8s Runner (schoenkitchen NS):** gitea-runner StatefulSet (gleiche Architektur) — RUNNING, declare erfolgreich (Sep 2026).
|
||||||
- **KNOWN BUG (Gitea 1.27.0):** `actions_ready_job` queue doesn't create `action_task` records for K8s-based runners via gRPC Declare. Only runner 73 (external, pre-registered) receives tasks. Fix: upgrade Gitea to 1.27.1+ or 1.28.x.
|
- **KNOWN BUG (Gitea 1.27.0):** `actions_ready_job` queue doesn't create `action_task` records for K8s-based runners via gRPC Declare. Only runner 73 (external, pre-registered) receives tasks. Fix: upgrade Gitea to 1.27.1+ or 1.28.x.
|
||||||
|
- **WARNUNG (Sep 2026):** `DEFAULT_ACTIONS_URL = https://gitea.com` ist deprecated in Gitea 1.27.0 — fallback zu `github`. Helm Values setzen nur `ENABLED: true`; der Wert kommt vom Chart-Default. Zu fixen via Helm values override: `DEFAULT_ACTIONS_URL: self`.
|
||||||
- **No runner management API** in Gitea 1.27.0 (`/api/v1/admin/runners` → 404). Manage via DB: `gitea.action_runner` table, `is_disabled` column, `deleted` for soft-delete.
|
- **No runner management API** in Gitea 1.27.0 (`/api/v1/admin/runners` → 404). Manage via DB: `gitea.action_runner` table, `is_disabled` column, `deleted` for soft-delete.
|
||||||
- Workflow: `.gitea/workflows/rebuild-infrastructure.yml`
|
- Workflow: `.gitea/workflows/rebuild-infrastructure.yml`
|
||||||
- Achtung: `actions/checkout@v4` versucht github.com zu erreichen — muss gemapped werden
|
- Achtung: `actions/checkout@v4` versucht github.com zu erreichen — muss gemapped werden
|
||||||
@@ -51,6 +54,19 @@ modified: "2026-07-26"
|
|||||||
- **Verwendung:** `GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o StrictHostKeyChecking=no" git push ssh://git@10.0.30.202:22/dominik/iac-homelab.git HEAD:main`
|
- **Verwendung:** `GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o StrictHostKeyChecking=no" git push ssh://git@10.0.30.202:22/dominik/iac-homelab.git HEAD:main`
|
||||||
- **Wichtig:** Gitea HTTP (Port 3000) hat keine externe IngressRoute — SSH (Port 22 via LoadBalancer) ist der einzige zuverlässige Weg für Git-Pushes von außerhalb K8s.
|
- **Wichtig:** Gitea HTTP (Port 3000) hat keine externe IngressRoute — SSH (Port 22 via LoadBalancer) ist der einzige zuverlässige Weg für Git-Pushes von außerhalb K8s.
|
||||||
|
|
||||||
|
## Known ArgoCD Issues (Sep 2026)
|
||||||
|
- **`gitea` app:** Sync=Succeeded, Health=Progressing — Ingress hat keine LoadBalancer IP (Traefik reporting gap). Kosmetisch, Funktionalität OK.
|
||||||
|
- **`gitea-config` app:** OutOfSync — `Deployment gitea-runner` driftet (manuelle Änderung am Image?). Sync würde helfen.
|
||||||
|
- **`immich` + `immich-config` apps:** OutOfSync — zwei Apps zeigen auf überlappende Pfade (`rendered` vs roh) im iac-homelab Repo. Architektur-Issue, kein Gitea-Problem.
|
||||||
|
- **`authelia` app:** Health=Progressing — gleiche Ingress-LB-IP Thematik wie gitea.
|
||||||
|
|
||||||
|
## Incident 2026-09-27: CSI RBAC → Gitea Pod Init-Crash
|
||||||
|
- **Symptom:** Gitea Pod 4+ Tage in `Init:0/3`, schoenkitchen runner in `PodInitializing`.
|
||||||
|
- **Root Cause:** Ceph CSI RBD provisioner `csi-attacher` hatte keine Berechtigung, `csinodes` zu lesen → VolumeAttachment blieb pending → PVC konnte nicht mounten.
|
||||||
|
- **Fix:** RBAC ClusterRole `ceph-rbd-external-attacher-runner` um `storage.k8s.io/csinodes` Resource ergänzt, provisioner Pods neu gestartet.
|
||||||
|
- Nach Fix: Alle 24 VolumeAttachments `true`, Gitea Pod Running nach Pod-Delete.
|
||||||
|
- **Lesson:** Bei CSI-basierten PVs immer zuerst VolumeAttachment Status prüfen, nicht nur PVC Phase.
|
||||||
|
|
||||||
## Related
|
## Related
|
||||||
- [[systems/rke2-kubernetes]]
|
- [[systems/rke2-kubernetes]]
|
||||||
- [[concepts/gitops-workflow]]
|
- [[concepts/gitops-workflow]]
|
||||||
|
|||||||
@@ -10,8 +10,10 @@ modified: "2026-07-24"
|
|||||||
|
|
||||||
## Deployment
|
## Deployment
|
||||||
- **Namespace:** hindsight (K8s)
|
- **Namespace:** hindsight (K8s)
|
||||||
- **API:** LoadBalancer `10.0.30.201:9177` (**NICHT localhost**)
|
- **API:** LoadBalancer `10.0.30.208:9177` (**NICHT localhost**, NICHT .201 — dort sitzt seit Rebuild 2026-08-01 Grafana!)
|
||||||
- **Health:** `curl -s http://10.0.30.201:9177/health`
|
- **Health:** `curl -s http://10.0.30.208:9177/health`
|
||||||
|
- **Port-Mapping:** Svc 9177 → Container 8888; NodePort-Fallback 31577 (jeder Node)
|
||||||
|
- **LB-IP-Drift-Warnung:** LB-IPs verschoben sich beim Rebuild 2026-08-01 (hindsight .201→.208). Bei Health-Failure IMMER zuerst `kubectl -n hindsight get svc hindsight-api` gegenprüfen, statt Referenz-IP zu vertrauen
|
||||||
- **Backend:** PostgreSQL + pgvector
|
- **Backend:** PostgreSQL + pgvector
|
||||||
- **Config:** `~/.hermes/hindsight/config.json` (api_url gesetzt)
|
- **Config:** `~/.hermes/hindsight/config.json` (api_url gesetzt)
|
||||||
|
|
||||||
@@ -32,6 +34,9 @@ modified: "2026-07-24"
|
|||||||
- **retain_every_n_turns: 2** — reduziert Duplikate
|
- **retain_every_n_turns: 2** — reduziert Duplikate
|
||||||
- Vor Batch-Retain: Health-Check, dann erst retain calls feuern
|
- Vor Batch-Retain: Health-Check, dann erst retain calls feuern
|
||||||
- Meta-Memory Cleanup Cron: 1st of Month 03:00 (3 Dedup Rounds, Archive >6 Monate)
|
- Meta-Memory Cleanup Cron: 1st of Month 03:00 (3 Dedup Rounds, Archive >6 Monate)
|
||||||
|
- Decay & Importance Scoring: 1st of Month 03:30 (score=access×age_decay, archive score<0.15)
|
||||||
|
- User Drift Detection: Mondays 04:00 (compare 14d vs 90d baseline)
|
||||||
|
- Skill Health Check: Mondays 04:30 (scan 187 SKILL.md for issues)
|
||||||
- Nightly Dream Cycle + Weekly Insight Digest
|
- Nightly Dream Cycle + Weekly Insight Digest
|
||||||
|
|
||||||
## What NOT to store in Hindsight
|
## What NOT to store in Hindsight
|
||||||
|
|||||||
@@ -0,0 +1,70 @@
|
|||||||
|
---
|
||||||
|
title: "Home Assistant"
|
||||||
|
category: systems
|
||||||
|
tags: [homeassistant, smart-home, mqtt, automation]
|
||||||
|
created: "2026-09-27"
|
||||||
|
modified: "2026-09-27"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Home Assistant
|
||||||
|
|
||||||
|
> Smart Home Zentrale auf Proxmox. Steuerung, Automatisierung und Benachrichtigungen.
|
||||||
|
|
||||||
|
## Infrastruktur
|
||||||
|
- **Host:** `10.0.30.10` (HAOS VM)
|
||||||
|
- **SSH:** `hassio@10.0.30.10` (PW: 1P Vault "Hermes")
|
||||||
|
- **URL:** `https://homeassistant.familie-schoen.com`
|
||||||
|
- **Container:** `homeassistant` (Docker)
|
||||||
|
|
||||||
|
## MQTT
|
||||||
|
- **Broker:** Mosquitto Add-on v7.1.1 auf localhost:1883
|
||||||
|
- **HA MQTT User:** ehemals `opendtu` (BROKEN — ACL blockiert), gefixt
|
||||||
|
- **ACL:** `/etc/mosquitto/acl` definiert `user homeassistant` + `user addons`
|
||||||
|
- **Auth Plugin:** `go-auth.so` (files,http backends)
|
||||||
|
|
||||||
|
## Automations
|
||||||
|
- **File:** `/config/automations.yaml` (35 Automations)
|
||||||
|
- **Schreibmethode:** SSH → `docker exec homeassistant chmod 666`, danach restore 644
|
||||||
|
- **Reload:** REST API mit JWT (HS256, signed from `/config/.storage/auth`)
|
||||||
|
|
||||||
|
### Frigate Einfahrt Notification (ID 1714065780832)
|
||||||
|
- MQTT trigger `frigate/events` → filter `type=='new'` + cameras [Einfahrt,Terrasse] + labels [person,car,...]
|
||||||
|
- Mode: `single`, cooldown 600s
|
||||||
|
- Dedup: `input_text.frigate_last_event_id`
|
||||||
|
- Flow: delay 10s → snapshot → notify iPhone → optional LLM Vision
|
||||||
|
- Siehe [[systems/frigate]]
|
||||||
|
|
||||||
|
## Notify Services
|
||||||
|
| Service | Status |
|
||||||
|
|---------|--------|
|
||||||
|
| `notify.mobile_app_iphone_dominik` | ✅ aktiv |
|
||||||
|
| `notify.mobile_app_sarahs_iphone_app` | verfügbar |
|
||||||
|
| `notify.mobile_app_ipad_2` | verfügbar |
|
||||||
|
| `notify.mobile_app_sm_x205` | verfügbar |
|
||||||
|
|
||||||
|
## Entitäten
|
||||||
|
- Kameras: `camera.einfahrt_2`, `camera.terrasse_2` (Suffix `_2` wegen verwaister Integration)
|
||||||
|
- Motion: `binary_sensor.einfahrt_motion_2`, `binary_sensor.terrasse_motion_2`
|
||||||
|
- Input Text: `input_text.frigate_last_event_id` (dedup storage, max 255 chars)
|
||||||
|
|
||||||
|
## Snapshots
|
||||||
|
- Gespeichert: `/config/www/snapshots/{camera}_latest.jpg`
|
||||||
|
- URL: `/local/snapshots/` (HTTP 200, keine Auth)
|
||||||
|
|
||||||
|
## JWT Auth
|
||||||
|
1. `jwt_key` aus `/config/.storage/auth` lesen
|
||||||
|
2. Client `Hermes_202606` (token id `1444b2c6757b4d66a6f8f8e5899b4e6a`)
|
||||||
|
3. HS256 signieren, Bearer Header
|
||||||
|
4. `?return_response=true` für Service-Call Responses
|
||||||
|
|
||||||
|
## Integrations
|
||||||
|
- Frigate (MQTT)
|
||||||
|
- LLM Vision (noris AI `gemma-4-31b-it`)
|
||||||
|
- Tibber (Strom)
|
||||||
|
- Marstek Speicher (VENUS-E, IP 10.0.50.113)
|
||||||
|
- Various sensors (Xiaomi BLE, etc.)
|
||||||
|
|
||||||
|
## Related
|
||||||
|
- [[systems/frigate]] — NVR Integration
|
||||||
|
- [[systems/noris-ai]] — LLM Vision Provider
|
||||||
|
- [[reference/ip-map]] — IP Assignments
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
# Laya Decision Model
|
||||||
|
|
||||||
|
> 421M param ModernBERT-large Classifier (Apache 2.0) auf CT152, CPU-only.
|
||||||
|
> Nutze in normalen Sessions fuer schnelle Klassifizierung, Binaerentscheidungen und Pre-Filter.
|
||||||
|
|
||||||
|
## Zugang
|
||||||
|
- **Endpoint:** `POST http://10.0.30.152:8000/predict`
|
||||||
|
- **Payload:** `{"state": "<text>", "questions": {...}}`
|
||||||
|
- **Health:** `GET http://10.0.30.152:8000/health`
|
||||||
|
- **Presets:** `GET http://10.0.30.152:8000/presets/{router|guard|moderation|triage}`
|
||||||
|
|
||||||
|
## Entscheidungstypen (Primitives)
|
||||||
|
|
||||||
|
| Typ |用途 | Return Fields |
|
||||||
|
|-----|------|---------------|
|
||||||
|
| `choice` | Klassifizierung in N Labels | `choice`, `answer_confidence`, `probabilities` |
|
||||||
|
| `noul` | Ja/Nein mit Wahrscheinlichkeit | `noul` (0..1), `answer_confidence` |
|
||||||
|
| `score` | Ordinale Bewertung | `score`, `answer_confidence` |
|
||||||
|
|
||||||
|
## Verwendung in Sessions
|
||||||
|
|
||||||
|
**Praeferieren fuer:**
|
||||||
|
- Pre-Filter vor teuren LLM-Calls (z.B. "ist diese Email eine Rechnung?" → nur bei "ja" LLM aufrufen)
|
||||||
|
- Binaerentscheidungen: alert/skip, escalate/ignore, move/keep
|
||||||
|
- Multi-Kategorie-Klassifizierung mit Confidence
|
||||||
|
- Gatekeeping: Notification-Suppression, Alert-Filtering
|
||||||
|
|
||||||
|
**NICHT geeignet fuer:**
|
||||||
|
- Textgenerierung / Zusammenfassungen (dafür LLM verwenden)
|
||||||
|
- Komplexe Reasoning-Tasks
|
||||||
|
- Embeddings / Semantische Suche (dafür Harrier/Hindsight)
|
||||||
|
|
||||||
|
## Example Call
|
||||||
|
|
||||||
|
```python
|
||||||
|
import json, urllib.request
|
||||||
|
|
||||||
|
payload = json.dumps({
|
||||||
|
"state": "Von: amazon.de\nBetreff: Bestellbestätigung #12345",
|
||||||
|
"questions": {
|
||||||
|
"kategorie": {
|
||||||
|
"type": "choice",
|
||||||
|
"instructions": "Welche Kategorie?",
|
||||||
|
"criteria": {
|
||||||
|
"rechnung": "Rechnung, Invoice",
|
||||||
|
"bestellung": "Bestellbestätigung, Order",
|
||||||
|
"werbung": "Newsletter, Marketing"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"ignorieren": {
|
||||||
|
"type": "noul",
|
||||||
|
"instructions": "Soll ignoriert werden?"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}).encode()
|
||||||
|
|
||||||
|
req = urllib.request.Request(
|
||||||
|
"http://10.0.30.152:8000/predict",
|
||||||
|
data=payload,
|
||||||
|
headers={"Content-Type": "application/json"}
|
||||||
|
)
|
||||||
|
result = json.loads(urllib.request.urlopen(req, timeout=30).read())
|
||||||
|
# result["answers"]["kategorie"]["choice"] → "bestellung"
|
||||||
|
# result["answers"]["kategorie"]["answer_confidence"] → 0.99
|
||||||
|
```
|
||||||
|
|
||||||
|
## Performance
|
||||||
|
- Latenz: ~1.5-1.7s warm cache (CPU)
|
||||||
|
- Load time: ~18s (cold start)
|
||||||
|
- systemd service: `laya.service` (enabled, onboot)
|
||||||
|
|
||||||
|
## Einsatzgebiete (aktiv)
|
||||||
|
- **Rechnungen-Organizer** (Cron `f773f8c23230`): 12-Kategorie Email-Klassifizierung, 3-Schichten-Safety
|
||||||
|
- Weitere Kandidaten: Beleg-Sammler, SRE Network Recon, Backup Digest, Frigate Event Gate
|
||||||
|
|
||||||
|
## Constraints
|
||||||
|
- 421M Modell → niedrige Confidence bei ambiguous Inputs (Feature, nicht Bug!)
|
||||||
|
- Batch funktioniert nicht — ein Request pro Input-Instanz
|
||||||
|
- `noul` ist nicht zuverlaessig fuer kritische Entscheidungen allein — immer mit `choice` kombinieren
|
||||||
|
- Confidence-Threshold empfohlen (≥0.6 fuer Moves, ≥0.5 fuer Ignores)
|
||||||
|
|
||||||
|
## Related
|
||||||
|
- [[concepts/email-organization]] — Rechnungs-Organizer Architecture
|
||||||
|
- [[entities/infrastructure]] — CT152 auf proxmox6
|
||||||
|
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
---
|
||||||
|
title: "noris AI Platform"
|
||||||
|
category: systems
|
||||||
|
tags: [ai, llm, noris, gpu, embeddings]
|
||||||
|
created: "2026-09-27"
|
||||||
|
modified: "2026-09-29"
|
||||||
|
---
|
||||||
|
|
||||||
|
# noris AI Platform (ai.noris.de)
|
||||||
|
|
||||||
|
> Interne AI-Plattform der noris Network AG. Bereitstellung von LLMs, Embeddings und Image Generation.
|
||||||
|
|
||||||
|
## Endpoints
|
||||||
|
- **Chat:** `https://ai.noris.de/v1/chat/completions`
|
||||||
|
- **Embeddings:** `https://ai.noris.de/v1/embeddings`
|
||||||
|
- **Images:** ⚠️ `/v1/images/generations` wird vom Bifrost Gateway **NICHT** unterstützt. Image-Gen-Modelle (qwen-image-2-1) werden über `/v1/chat/completions` angesprochen — das Bild kommt als base64-PNG im `content`-Array zurück (Typ `image_url`, `data:image/png;base64,...`). Siehe `references/vllm-image-generation.md` im Skill `serving-llms-vllm`.
|
||||||
|
|
||||||
|
## Modelle
|
||||||
|
| Typ | Modell-ID | Hinweise |
|
||||||
|
|-----|-----------|----------|
|
||||||
|
| Flagship LLM | `glm-5-2` | Primary, OpenRouter-kompatibel |
|
||||||
|
| General | `gemma-4-31b-it` | Vision-fähig, genutzt von LLM Vision |
|
||||||
|
| Large MoE | `gpt-oss-120b` | |
|
||||||
|
| Mid-range | `qwen3.6-27b` | |
|
||||||
|
| Mid-range | `qwen3.8-27b` | |
|
||||||
|
| Fast | `ds-v4-flash` | Low-latency, Paperless OCR |
|
||||||
|
| Embedding | `harrier` | Vektorembeddings |
|
||||||
|
| Image Gen | `qwen-image-2-1` | Via `/v1/chat/completions` (NOT images/generations). Base64-PNG im content-Array. ~30s/ Bild. |
|
||||||
|
|
||||||
|
## Verbraucher
|
||||||
|
- **Hermes Agent** — Primärmodell `glm-5-2` via OpenRouter
|
||||||
|
- **HA LLM Vision** — `gemma-4-31b-it` für Bildanalyse (Frigate Events)
|
||||||
|
- **Paperless** — `ds-v4-flash` für OCR/Kategorisierung
|
||||||
|
- **Personal Coach Bot** — `glm-5-2` via noris direkt
|
||||||
|
- **Dynamic Coach** — `glm-5-2` via noris direkt, `qwen-image-2-1` für Visualisierungen
|
||||||
|
|
||||||
|
## Related
|
||||||
|
- [[systems/frigate]] — nutzt noris AI für Event-Klassifizierung
|
||||||
|
- [[systems/homeassistant]] — LLM Vision Integration
|
||||||
|
- [[systems/paperless]] — OCR via ds-v4-flash
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
---
|
||||||
|
title: "Paperless-ngx"
|
||||||
|
category: systems
|
||||||
|
tags: [paperless, documents, oidc, ocr]
|
||||||
|
created: "2026-09-27"
|
||||||
|
modified: "2026-09-27"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Paperless-ngx
|
||||||
|
|
||||||
|
> Dokumentenmanagement mit OCR, OIDC-Login und AI-Kategorisierung.
|
||||||
|
|
||||||
|
## Zugriff
|
||||||
|
- **URL:** `https://dokumente.familie-schoen.com`
|
||||||
|
- **mTLS:** `https://dokumente-mtls.familie-schoen.com` (auto-login als `dominik`)
|
||||||
|
- **PKCS12:** `dominik-dokumente-mtls.p12` (PW: siehe 1P Vault "Hermes")
|
||||||
|
|
||||||
|
## Auth
|
||||||
|
- OIDC via Authelia (`dominik@schoen.eu`)
|
||||||
|
- Break-Glass lokaler User: `dominik`
|
||||||
|
- mTLS Client Cert: CN=dominik, gültig bis Juli 2028
|
||||||
|
|
||||||
|
## Konfiguration
|
||||||
|
- **AI Backend:** `ds-v4-flash@ai.noris.de` (noris AI)
|
||||||
|
- **Memory:** 4Gi (PAT-002)
|
||||||
|
- **Mail Import:** `dokumente@familie-schoen.com` (iCloud mailbox, max 30 Tage)
|
||||||
|
- **Owner:** dominik
|
||||||
|
|
||||||
|
## Known Issue
|
||||||
|
- PAT-002: Memory-Limit 4Gi erforderlich, sonst OOM bei großen OCR-Batches
|
||||||
|
|
||||||
|
## Related
|
||||||
|
- [[concepts/credential-policy]] — 1Password, mTLS Zertifikate
|
||||||
|
- [[systems/noris-ai]] — ds-v4-flash für OCR
|
||||||
@@ -3,14 +3,14 @@ title: Proxmox VE Cluster
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [proxmox, virtualization, lxc, qemu, pve]
|
tags: [proxmox, virtualization, lxc, qemu, pve]
|
||||||
created: "2026-04-28"
|
created: "2026-04-28"
|
||||||
modified: "2026-07-24"
|
modified: "2026-09-30"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Proxmox VE Cluster
|
# Proxmox VE Cluster
|
||||||
|
|
||||||
## Cluster-Konfiguration
|
## Cluster-Konfiguration
|
||||||
- **Version:** PVE 9.2.3
|
- **Version:** PVE 9.2.20, Kernel 7.0.14-19-pve (upgraded 2026-09-25)
|
||||||
- **Nodes:** 9 (Quorum OK)
|
- **Nodes:** 8 (Quorum OK, proxmox2 dauerhaft entfernt)
|
||||||
- **Hypervisoren:** 10.0.20.x
|
- **Hypervisoren:** 10.0.20.x
|
||||||
- **Guests:** ~30 LXC + ~10 QEMU VMs
|
- **Guests:** ~30 LXC + ~10 QEMU VMs
|
||||||
|
|
||||||
@@ -50,13 +50,92 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
|||||||
- Benötigte modprobe.d Config:
|
- Benötigte modprobe.d Config:
|
||||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||||
- Worker-05 (VM 102) läuft auf ms-a2-2 mit funktionierendem GPU-Passthrough
|
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
|
||||||
|
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
|
||||||
|
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
|
||||||
|
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
|
||||||
|
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
|
||||||
|
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
|
||||||
|
|
||||||
## Bekannte Probleme
|
## Bekannte Probleme
|
||||||
- osd.0 NVMe 92% Wear — Austausch planen
|
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||||
- osd.2/5 nearfull (93-94%) — entlasten
|
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
||||||
- CT110 kaputte libc — Reparatur ausstehend
|
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
||||||
- ms-a2-1 GPU-Passthrough: Config gefixt, Reboot zur Verifikation ausstehend
|
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
|
||||||
|
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
|
||||||
|
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
|
||||||
|
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
|
||||||
|
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
|
||||||
|
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
|
||||||
|
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
|
||||||
|
|
||||||
|
## HA Rules (PVE 9.2 Rules System)
|
||||||
|
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||||
|
Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
||||||
|
|
||||||
|
**Resource-Affinity (negative = anti-collocation):**
|
||||||
|
|
||||||
|
*RKE2 Control Plane (etcd-Quorum braucht 2/3):*
|
||||||
|
- `rke2-cp-anti-112-122`: vm:112 ↔ vm:122 — nie auf gleichem Host
|
||||||
|
- `rke2-cp-anti-112-126`: vm:112 ↔ vm:126 — nie auf gleichem Host
|
||||||
|
- `rke2-cp-anti-122-126`: vm:122 ↔ vm:126 — nie auf gleichem Host
|
||||||
|
|
||||||
|
*RKE2 Worker:*
|
||||||
|
- `rke2-worker-anti-102-128`: vm:102 ↔ vm:128 — nie auf gleichem Host
|
||||||
|
- `rke2-worker-anti-102-139`: vm:102 ↔ vm:139 — nie auf gleichem Host
|
||||||
|
- `rke2-worker-anti-128-139`: vm:128 ↔ vm:139 — nie auf gleichem Host
|
||||||
|
|
||||||
|
*Galera:*
|
||||||
|
- `galera-anti-300-301`: vm:300 ↔ vm:301 — nie auf gleichem Host
|
||||||
|
- `galera-anti-300-302`: vm:300 ↔ vm:302 — nie auf gleichem Host
|
||||||
|
- `galera-anti-301-302`: vm:301 ↔ vm:302 — nie auf gleichem Host
|
||||||
|
|
||||||
|
*MaxScale:*
|
||||||
|
- `maxscale-anti-310-311`: vm:310 ↔ vm:311 — nie auf gleichem Host
|
||||||
|
|
||||||
|
**Node-Affinity (non-strict, failover allowed):**
|
||||||
|
|
||||||
|
*RKE2 CP:*
|
||||||
|
- `na-vm112`: vm:112 → proxmox3, proxmox5, proxmox7
|
||||||
|
- `na-vm122`: vm:122 → ms-a2-1, ms-a2-2
|
||||||
|
- `na-vm126`: vm:126 → proxmox4, proxmox5, proxmox6
|
||||||
|
|
||||||
|
*RKE2 Worker:*
|
||||||
|
- `na-vm102`: vm:102 → proxmox5, proxmox6, proxmox7 (NICHT ms-a2-2!)
|
||||||
|
- `na-vm128`: vm:128 → ms-a2-2, proxmox5, proxmox7
|
||||||
|
- `na-vm139`: vm:139 → n5pro, proxmox3, proxmox4
|
||||||
|
|
||||||
|
*Hermes:*
|
||||||
|
- `na-vm230`: vm:230 → n5pro, proxmox5, proxmox6
|
||||||
|
|
||||||
|
*Galera/MaxScale:*
|
||||||
|
- `na-vm300`: vm:300 → n5pro, proxmox3, proxmox4
|
||||||
|
- `na-vm301`: vm:301 → proxmox6, proxmox5, proxmox7
|
||||||
|
- `na-vm302`: vm:302 → ms-a2-2, ms-a2-1
|
||||||
|
- `na-vm310`: vm:310 → proxmox7, proxmox5, proxmox4
|
||||||
|
- `na-vm311`: vm:311 → ms-a2-1, ms-a2-2
|
||||||
|
|
||||||
|
**Aktuelle Verteilung (alle Anti-Collocation erfüllt):**
|
||||||
|
| Role | VM | Node |
|
||||||
|
|------|----|------|
|
||||||
|
| RKE2 CP-01 | 112 | proxmox3 |
|
||||||
|
| RKE2 CP-02 | 122 | ms-a2-1 |
|
||||||
|
| RKE2 CP-03 | 126 | proxmox4 |
|
||||||
|
| RKE2 Worker-01 | 128 | proxmox5 |
|
||||||
|
| RKE2 Worker-04 | 139 | n5pro |
|
||||||
|
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
||||||
|
| Hermes-Agent-01 | 230 | n5pro |
|
||||||
|
| Galera db1 | 300 | n5pro |
|
||||||
|
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
|
||||||
|
| Galera db3 | 302 | ms-a2-2 |
|
||||||
|
| MaxScale-01 | 310 | proxmox7 |
|
||||||
|
| MaxScale-02 | 311 | ms-a2-1 |
|
||||||
|
|
||||||
|
> ⚠️ PVE 9.2 Constraints:
|
||||||
|
> - Resources in resource-affinity rules dürfen keine multi-priority node-affinity haben (gleiche Priorität für alle Nodes erforderlich).
|
||||||
|
> - `ha-manager add` MUSS vor `ha-manager rules add` kommen — sonst "cannot use unmanaged resource".
|
||||||
|
> - Bei gleichzeitigem HA-Add + Anti-Collocation-Violation kann HA-Manager deadlocks (beide VMs auf `migrate` fest). Lösung: eine VM temporär aus HA entfernen, manuell migrieren, dann re-add.
|
||||||
|
> - Online-Migration von VMs mit hohen Memory-Writes (>12GB dirty pages) kann `broken pipe` fehlschlagen. Offline-Migration (stop→migrate→start) als Fallback.
|
||||||
|
|
||||||
## Related
|
## Related
|
||||||
- [[systems/ceph-cluster]]
|
- [[systems/ceph-cluster]]
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ title: RKE2 Kubernetes Cluster
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [kubernetes, rke2, cilium, argocd, cnpg, gitops]
|
tags: [kubernetes, rke2, cilium, argocd, cnpg, gitops]
|
||||||
created: "2026-07-24"
|
created: "2026-07-24"
|
||||||
modified: "2026-07-24"
|
modified: "2026-09-17"
|
||||||
---
|
---
|
||||||
|
|
||||||
# RKE2 Kubernetes Cluster
|
# RKE2 Kubernetes Cluster
|
||||||
@@ -22,16 +22,21 @@ modified: "2026-07-24"
|
|||||||
| cp-03 | 10.0.30.53 | Control Plane |
|
| cp-03 | 10.0.30.53 | Control Plane |
|
||||||
| worker-01 | 10.0.30.63 | Worker |
|
| worker-01 | 10.0.30.63 | Worker |
|
||||||
| worker-04 | 10.0.30.64 | Worker |
|
| worker-04 | 10.0.30.64 | Worker |
|
||||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2; 30.09. OOM-Zombie-Freeze behoben, siehe PAT-010) |
|
||||||
|
|
||||||
## Storage
|
## Storage
|
||||||
- **Ceph CSI**: ceph-flash (fast), ceph-hdd (bulk)
|
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
||||||
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`
|
- **⚠️ StorageClass IaC (2026-09-29):** `ceph-flash` + `ceph-hdd-replica` waren MANUELL erstellt (nicht in Git). Jetzt committed unter `clusters/main/storage/`. Alle RBD SCs MÜSSEN `controller-expand-secret-name/namespace` haben — fehlt dies, schlagen Volume Expansions fehl ("provided secret is empty") und triggern Retry-Loops die API Server überlasten.
|
||||||
|
- **⚠️ PVs snapshotten SC-Parameter zur Provisionierungszeit** — SC fixen reicht NICHT; bestehende PVs brauchen individuellen Patch mit `controllerExpandSecretRef` falls sie ohne erstellt wurden.
|
||||||
|
- **⚠️ StorageClass `parameters` sind IMMUTABLE** — können nicht gepatched werden, müssen gelöscht+neu erstellt werden.
|
||||||
|
- **⚠️ Snapshot-CRD-Flavor (seit Rebuild 2026-08-01):** `volumesnapshotclasses`-CRD (RKE2-Addon `rke2-snapshot-controller-crd`) hat **FLAT-Schema** — `driver`/`deletionPolicy`/`parameters` auf TOP-LEVEL, `spec:` existiert nicht im Schema (Upstream-CRD wäre nested!). Niemals struktur-intuitiv "reparieren" — CRD-Schema lesen. Live-VSCs: `ceph-rbd-snapclass` (default) + `cephfs-snapclass`
|
||||||
|
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`; Specs **100Gi data / 20Gi WAL** (grow-only — Shrink wird von CNPG-Admission verboten, Git immer nach oben alignieren)
|
||||||
|
- **⚠️ etcd auf Ceph RBD (~20ms WAL fsync auf allen CP-Nodes)** — architektonisches Risiko; unter API Server Write Pressure → gRPC DeadlineExceeded Cascade. Lokales NVMe für etcd data dirs empfohlen.
|
||||||
|
|
||||||
## GitOps
|
## GitOps
|
||||||
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
|
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
|
||||||
- 15 Applications via SSH (`ssh://git@10.0.30.202:22/dominik/iac-homelab.git`)
|
- **Feed-Quelle (Stand 2026-09-17):** `http://10.0.30.105:3000/...` = Legacy-CT108-Gitea (noch aktiv!). Geplanter Flip auf `ssh://git@10.0.30.200:22/dominik/iac-homelab.git` (end-to-end verifiziert), danach CT108-Stop (Freigabe Dominik)
|
||||||
- **Gitea SSH user is `git`, not `gitea`** — `git.schoen.codes` resolves to Traefik (10.0.30.203, HTTP only), SSH is on 10.0.30.202:22
|
- **Gitea SSH user is `git`, not `gitea`** — SSH-LoadBalancer = **10.0.30.200:22** (svc gitea-ssh, targetPort 2222, NodePort 31441). Legacy-.202 ist TOD (Timeout). Repo-Historien am 2026-09-17 konsolidiert (identische Tips auf beiden Remotes, Merge `d76d3e8` + `8c741c5`)
|
||||||
- Workflow: IaC repo auschecken → ändern → commit+push → ArgoCD sync → verify → lokale Kopie löschen
|
- Workflow: IaC repo auschecken → ändern → commit+push → ArgoCD sync → verify → lokale Kopie löschen
|
||||||
- Siehe [[concepts/gitops-workflow]]
|
- Siehe [[concepts/gitops-workflow]]
|
||||||
|
|
||||||
@@ -61,6 +66,19 @@ modified: "2026-07-24"
|
|||||||
- IaC: unified `amd-gpu` PCI mapping (n5pro + ms-a2-1 + ms-a2-2) + dynamic hostpci (one worker block, gpu flag)
|
- IaC: unified `amd-gpu` PCI mapping (n5pro + ms-a2-1 + ms-a2-2) + dynamic hostpci (one worker block, gpu flag)
|
||||||
- ms-a2-1 has kernel BUG with 1002:13c0 (renderD128 missing) — worker-05 moved to ms-a2-2 (identical hardware)
|
- ms-a2-1 has kernel BUG with 1002:13c0 (renderD128 missing) — worker-05 moved to ms-a2-2 (identical hardware)
|
||||||
|
|
||||||
|
## Capacity Management
|
||||||
|
- **Descheduler** v0.30.0 deployed als ArgoCD-App (`clusters/main/descheduler/manifests.yaml`)
|
||||||
|
- Plugin: `LowNodeUtilization` (thresholds 30/50/20, targetThresholds 70/70/60)
|
||||||
|
- Interval: 2m (`--descheduling-interval=2m`)
|
||||||
|
- `nodeFit: true` auf Profile-Ebene (nicht Plugin-Ebene — v0.30 Schema!)
|
||||||
|
- RBAC: benötigt `list/watch` auf `namespaces` (fehlt in Default-Helm-Chart)
|
||||||
|
- Raw Manifests statt Helm Chart (Chart 0.30 hat params/args-Bug, Chart 0.35 hat falsche API-Version)
|
||||||
|
- Siehe Solution Doc `2026-09-30-k8s-cp01-relief-descheduler.md`
|
||||||
|
- **CP-01 Relief (2026-09-30):** hindsight-api + hindsight-postgres + paperless von cp-01 → worker-05 migriert
|
||||||
|
- cp-01 RAM: 93% → 43%; worker-05 bei ~28%
|
||||||
|
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren
|
||||||
|
- Live-only nodeSelector (hindsight-api, Sep-17-Hotfix) war nie in Git → Live-Patch nötig
|
||||||
|
|
||||||
## Known Pitfalls
|
## Known Pitfalls
|
||||||
- `enableServiceLinks: false` bei Apps deren Service-Name mit Env-Vars kollidiert (z.B. Paperless `PAPERLESS_PORT`)
|
- `enableServiceLinks: false` bei Apps deren Service-Name mit Env-Vars kollidiert (z.B. Paperless `PAPERLESS_PORT`)
|
||||||
- ArgoCD `--force` kann nicht mit ServerSideApply kombiniert werden
|
- ArgoCD `--force` kann nicht mit ServerSideApply kombiniert werden
|
||||||
|
|||||||
@@ -0,0 +1,74 @@
|
|||||||
|
---
|
||||||
|
title: Sarah-Hermes (VM107)
|
||||||
|
category: systems
|
||||||
|
tags: [hermes, webui, vm, sarah, telegram, backup]
|
||||||
|
created: "2026-09-30"
|
||||||
|
modified: "2026-09-30"
|
||||||
|
related: [systems/rke2-kubernetes, reference/ip-map, systems/noris-ai]
|
||||||
|
---
|
||||||
|
|
||||||
|
# Sarah-Hermes (VM107)
|
||||||
|
|
||||||
|
Zweite, vollständig isolierte Hermes-Instanz für Sarah (Allround-Assistentin:
|
||||||
|
Erinnerungen, Planung, Smalltalk). Getrennte Memories/Sessions/Skills/Keys von
|
||||||
|
den 5 Owner-Profilen.
|
||||||
|
|
||||||
|
## Eckdaten
|
||||||
|
|
||||||
|
| Attribut | Wert |
|
||||||
|
|----------|------|
|
||||||
|
| VM-ID | 107 (auto-allokiert) |
|
||||||
|
| Node | n5pro (Template 9000 debian-12-cloudinit) |
|
||||||
|
| IP | 10.0.30.66/24 (static via DHCP reservation) |
|
||||||
|
| Specs | 2 vCPU / 4 GB RAM / 32 GB Disk (vm_disks/RBD) |
|
||||||
|
| Access | SSH `debian@10.0.30.66` mit `~/.ssh/id_ed25519_cloudinit` |
|
||||||
|
| Tofu | `iac-homelab/epic-8-sarah-hermes/tofu/` |
|
||||||
|
| Ansible | `iac-homelab/epic-8-sarah-hermes/ansible/` |
|
||||||
|
| Plan | `iac-homelab/docs/plans/2026-09-30-sarah-hermes-vm.md` |
|
||||||
|
|
||||||
|
## Stack
|
||||||
|
|
||||||
|
- **hermes-webui** (ghcr.io/nesquena/hermes-webui:latest), Single-Container,
|
||||||
|
Port 8787 (0.0.0.0 gebunden, UFW erlaubt nur LAN), Password-Auth.
|
||||||
|
Compose: `/home/debian/hermes-webui/docker-compose.yml` auf der VM.
|
||||||
|
- **Agent-Runtime:** nousresearch/hermes-agent geklont nach
|
||||||
|
`~/.hermes/hermes-agent` auf der VM; installiert in `/app/venv` im Container
|
||||||
|
(editable). **Wichtig:** Core-Deps im pyproject sind hinter
|
||||||
|
`python_version >= '3.14'`-Markern gepinnt → auf Python 3.12 installiert
|
||||||
|
`-e .` NULL Deps. Manual-Dep-Bootstrapping nötig (siehe Pitfalls).
|
||||||
|
- **LLM:** noris-Provider (`https://ai.noris.de/v1`, Default-Modell
|
||||||
|
`vllm/release/glm-5-2`), Key via `HERMES_CUSTOM_NORIS_API_KEY` aus
|
||||||
|
`.env` (Compose mapped explizit in den Container).
|
||||||
|
- **Telegram:** geplant (blockiert auf BotFather-Token von Sarah/Dominik).
|
||||||
|
Bei Aktivierung: `gateway.telegram_enabled: true` + Token in `.env`;
|
||||||
|
Webhook/API-Server-Ports bleiben disabled (Konfliktvermeidung).
|
||||||
|
|
||||||
|
## Secrets (1Password, Vault: Hermes)
|
||||||
|
|
||||||
|
- `sarah-hermes-webui` → HERMES_WEBUI_PASSWORD
|
||||||
|
- `sarah-hermes-noris-key` → HERMES_CUSTOM_NORIS_API_KEY
|
||||||
|
|
||||||
|
## Backup
|
||||||
|
|
||||||
|
- Daily vzdump-Job `backup-2e8a34e3-66cb` (23:00, `all=1`, exclude 301,302,310,311,137,147,501)
|
||||||
|
→ **deckt VM107 ab** (Storage `noris_v4` = PBS Datastore `noris` @ 10.0.30.119).
|
||||||
|
- Manueller Verify-Lauf am 30.09.: TASK OK in 48s (inkrementell, 88% reuse).
|
||||||
|
- Offsite: folgt dem regulären `push-offsite` Sync-Job (Pull↔Push-Korrektur
|
||||||
|
vom 19.09.).
|
||||||
|
|
||||||
|
## Known Issues / Pitfalls
|
||||||
|
|
||||||
|
1. **Python-Version-Mismatch:** hermes-agent pyproject pins Core-Deps an
|
||||||
|
`python_version >= '3.14'`; WebUI-Container läuft auf 3.12 → `-e .`
|
||||||
|
installiert keine Deps. Fix: manuell `pip install` der gepinschten Pakete
|
||||||
|
+ iterativer Missing-Import-Loop bis `import run_agent` klappt.
|
||||||
|
2. **hermes update Ownership-Konflikt:** `hermes update`-Completion beschwert
|
||||||
|
sich über uid 0 vs. uid 1000 auf `/app/venv/bin/hermes-acp`. Kosmetisch,
|
||||||
|
Betriebsbetrieb unbeeinträchtigt. Fix-Idee: `chown` im Entry-Point.
|
||||||
|
3. **telegram-bridge Notify-Target defekt** (seit 19.09., Connection refused):
|
||||||
|
betrifft vzdump-Notifications clusterweit, nicht nur VM107. Separater Fix.
|
||||||
|
|
||||||
|
## Verification History
|
||||||
|
|
||||||
|
- 2026-09-30: Deploy + E2E-Test (Login 200, Agent-Antwort "HALLO" via glm-5-2).
|
||||||
|
- 2026-09-30: Manueller vzdump → TASK OK (48s, inkrementell).
|
||||||
Reference in New Issue
Block a user