Compare commits

...
23 Commits
Author SHA1 Message Date
Dominik Schön b527e4ff3c Auto-sync: 2026-09-30 2026-09-30 22:00:05 +00:00
Dominik Schön 3a57db957e log: ceph-monitoring-hardened Eintrag (mgr-union + SMART-Korrektur) 2026-09-30 18:23:53 +00:00
Dominik Schön cc2f237ca3 ceph: SMART-Audit rehabilitiert Kingston NVMe (kein Austausch); mgr-union scraping 2026-09-30 18:22:25 +00:00
Dominik Schön 531d6bfbdd docs(wiki): PAT-010 pve-oom-frozen-guest + proxmox6 drift fixes (worker-05 placement, oversubscription known-issue) 2026-09-30 14:15:18 +00:00
Dominik Schön a68665eca0 docs(wiki): sarah-hermes VM107 page + ip-map row + log entry 2026-09-30 13:49:08 +00:00
Dominik Schön d487b596d7 Auto-sync: 2026-09-29 2026-09-29 22:00:35 +00:00
Dominik Schön f09c6ef414 Auto-sync: 2026-09-28 2026-09-28 22:00:43 +00:00
Dominik Schön dfd3add11b Auto-sync: 2026-09-27 2026-09-27 22:00:24 +00:00
Dominik Schön d1464eb9e3 Restructure memory layers: USER.md/MEMORY.md cleanup, 4 new system pages, memory-layer-architecture concept 2026-09-27 15:15:18 +00:00
Dominik Schön 02f6a8e2b0 Auto-sync: 2026-09-26 2026-09-26 22:00:22 +00:00
Dominik Schön f3039765ed Auto-sync: 2026-09-21 2026-09-21 22:00:43 +00:00
Dominik Schön 1d036dbcdc Auto-sync: 2026-09-17 2026-09-17 22:00:39 +00:00
Dominik Schön fb47c2a182 Auto-sync: 2026-09-14 2026-09-14 22:00:45 +00:00
Dominik Schön b2e9512a0e Auto-sync: 2026-09-07 2026-09-07 22:00:23 +00:00
Dominik Schön 7c36a7600a Auto-sync: 2026-08-31 2026-08-31 22:00:17 +00:00
Dominik Schön 23e3702f0a chore: update skill-impact tracker with retrospective entry 2026-08-30 11:57:19 +00:00
Dominik Schön 5cf6a573f1 retro: compound-learning 30-day retrospective + 2 new patterns
New patterns:
- PAT-008: Ansible default_ipv6 fact missing on fresh VMs
- PAT-009: Stale NBD devices after RBD volume swap

Updated:
- index.md: added PAT-008, PAT-009
- log.md: retrospective entry with 4 solution docs + 2 patterns
- Solution docs dispatched to ~/docs/solutions/
2026-08-30 11:56:47 +00:00
Dominik Schön bbbe4f8985 feat: patterns/ directory + skill-impact tracker (WikiSkill-inspired)
- New patterns/ directory with 7 initial failure-mode patterns (PAT-001..007)
- skill-impact.md audit trail for skill modifications
- _template.md for future pattern creation
- index.md updated with Patterns section
- log.md entry for this change
- Inspired by arXiv:2608.27454 (WikiSkill)
2026-08-30 11:16:19 +00:00
Dominik Schön 6a7aed6f48 Auto-sync: 2026-08-24 2026-08-24 22:00:20 +00:00
Dominik Schön 5334f8ec07 Auto-sync: 2026-08-17 2026-08-17 22:00:52 +00:00
Dominik Schön 971edfea8e Auto-sync: 2026-08-13 2026-08-13 22:00:16 +00:00
Dominik Schön 222f286adc Auto-sync: 2026-08-12 2026-08-12 22:00:11 +00:00
Dominik Schön c392fc45f0 Auto-sync: 2026-07-31 2026-07-31 22:00:25 +00:00
37 changed files with 1785 additions and 114 deletions
+1 -1
View File
@@ -14,7 +14,7 @@ Infra-Changes **bevorzugt** über GitOps. Geht nicht immer. Ausnahmen müssen vo
## Workflow ## Workflow
1. IaC Repo lokal auschecken (`/root/iac-homelab` auf VM200) 1. IaC Repo lokal auschecken (`/root/iac-homelab` auf VM200)
2. Ändern (Tofu/Ansible/K8s Manifeste) 2. Ändern (Tofu/Ansible/K8s Manifeste)
3. Commit + Push (zu BEIDEN Gitea Remotes: origin + k8s) 3. Commit + Push zum EINEN kanonischen Remote: `ssh://git@10.0.30.202:22/dominik/iac-homelab.git` (Dual-Push-Habit führte 09/26 zur Ghost-Regression über CT108 — NIEMALS wieder zweites Remote pflegen)
4. ArgoCD sync (oder auto-sync) 4. ArgoCD sync (oder auto-sync)
5. Verify (kubectl get, curl, etc.) 5. Verify (kubectl get, curl, etc.)
6. Lokale Kopie löschen 6. Lokale Kopie löschen
+135
View File
@@ -0,0 +1,135 @@
---
title: "Memory Layer Architecture"
category: concepts
tags: [memory, architecture, hindsight, lcm, wiki, separation]
created: "2026-09-27"
modified: "2026-09-27"
---
# Memory Layer Architecture — Practical Rules
## Übersicht
7 Schichten mit klaren, nicht-überlappenden Rollen.
| # | Layer | System | Pfad/Ort | Rolle | Wann nutzen |
|---|-------|--------|----------|-------|-------------|
| L0 | Hot | MEMORY.md + USER.md | `~/.hermes/memories/` | System-Prompt Injection, komprimierte Facts | Jede Session (automatisch) |
| L1 | Curated | LLM Wiki | `~/.hermes/memory/` | Browsbares Wissen, git-synced | Wenn Details gebraucht werden |
| L2 | Semantic | Hindsight | K8s PostgreSQL | Vektor-Suche, semantisches Retrieval | Cross-Session Lookup |
| L3 | Procedural | Skills | `~/.hermes/skills/` | Wie-geht-es Prozeduren mit Pitfalls | Bei wiederkehrenden Tasks |
| L4 | Episodic | LCM + session_search | SQLite lcm.db / state.db | Rohe Gesprächsverläufe, Compaction | Innerhalb aktiver Session |
| L5 | Solutions | Solution Docs | `~/docs/solutions/` | Problem-Lösungs-Paare | Nach substanziellen Tasks |
| L6 | Project | AGENTS.md | pro Repo | Repo-spezifische Konventionen | Beim Arbeiten in einem Repo |
## Was wohin gehört — Entscheidungsbaum
```
Ist es ein Fact über Dominik persönlich?
→ USER.md (L0)
Ist es sein Kommunikationsstil, Safety-Rule, oder Workflow-Präferenz?
→ USER.md
Ist es ein persönlicher Umstand (Größe, Gewicht, Arbeitgeber)?
→ Hindsight (L2) — NICHT USER.md
Ist es ein Infra-Fakt (IP, Version, Konfiguration)?
→ LLM Wiki (L1) — systems/ oder reference/
Braucht er jede Session sofort?
→ MEMORY.md (L0) als 1-Zeiliger Pointer "→Wiki systems/xyz"
Ist es ein "Wie mache ich X"-Prozess?
→ Skills (L3)
Ist es ein "Wir hatten Problem Y, Lösung war Z"-Eintrag?
→ Solution Doc (L5) + Hindsight-Index (L2)
Ist es ein vergangenes Gesprächsereignis?
→ LCM (L4) kümmert sich automatisch darum
```
## L0 — MEMORY.md vs USER.md
### USER.md (1.375 chars max)
**Nur:** Dominiks Persönliche Präferenzen, Safety-Rules, Kommunikationsstil
- "NIEMALS ohne Approval löschen"
- "Action-first, kein Dom"
- "Crash+Alert statt Silent Skip"
- "Laya in Sessions nutzen"
**NICHT:** Infra-Fakten, IP-Adressen, Versionsnummern, Team-Struktur
### MEMORY.md (2.200 chars max)
**Nur:** Komprimierte Pointer auf Wiki-Seiten + nicht-wikiffähige Quick-Facts
- `Frigate v0.18 CT151. →Wiki systems/frigate`
- `Galera:VM300/301/302,VIP .70. →Wiki systems/galera-maxscale`
- Baufinanz-Zahlen (persönlich, nicht im Wiki)
- LinkedIn Persona (Workflow, nicht im Wiki)
**NICHT:** Vollständige Specs, Konfigurationsdetails, lange Beschreibungen
## L1 — LLM Wiki
### Wann erstellen/aktualisieren?
- Bei jeder infrastrukturellen Änderung (Compound-Learning Cycle Phase 4.7)
- Neue Seite wenn: neues System, neuer Service, neue architektonische Entscheidung
- Patch wenn: Versionswechsel, IP-Änderung, Known Issue hinzugekommen
### Wann NICHT?
- Procedures → Skills (L3)
- Problem-Solution Pairs → Solution Docs (L5)
- Credentials → 1Password
## L2 — Hindsight
### Was gehört rein?
- **Semantische Pointer** auf Solution Docs und wichtige Meilensteine
- **Cross-Session Kontext** der nicht im Wiki steht (z.B. "wir haben am Datum X entschieden...")
- **Episodisches Wissen** über Projekte, Entscheidungen, Learnings
### Was NICHT rein gehört?
- ❌ Infra-Topologie (→ L1 Wiki)
- ❌ Aktuelle IPs/Versionsnummern (→ L1 Wiki, werden schnell stale)
- ❌ Procedures (→ L3 Skills)
- ❌ Credentials (→ 1Password)
- ❌ Session-Fortschritt (→ L4 LCM)
### Bekanntes Problem: Hindsight Accumulation
Hindsight hat über Monate Infra-Fakten angesammelt die inzwischen teilweise
veraltet sind (9 vs 8 Nodes, alte OSD-Counts). Hindsight bietet keine
Delete-Funktion via Tools. Strategie:
1. Zukünftig nur noch semantische Pointer + Entscheidungen speichern
2. Topologie-Fakten nicht mehr via `hindsight_retain` ablegen
3. Veraltete Einträge ignorieren — Hindsight Ranking bevorzugt neuere Memories
## L4 — LCM (Local Conversation Memory)
### Rolle
- **Innerhalb aktiver Session:** Compaction, Summary-DAG, Fresh Tail
- **Cross-Session (lcm.db):** Rohe Nachrichten + Summary Nodes aller Sessions
- **Kein manuelles Management nötig** — läuft automatisch
### Wann manuell eingreifen?
- `lcm_status` bei Verdacht auf Compression-Problemen
- `lcm_grep` um exakte Phrasen in vergangenen Sessions zu finden
- `lcm_recall` für bedeutungsbasierte Cross-Session-Suche
- Niemals manuell Daten in LCM schreiben — es ist Auto-Managed
## L5 — Solution Docs
### Wann erstellen?
- Nach substanziellem Bugfix (≥5 Tool-Calls)
- Nach Architektur-Änderung
- Nach komplexer Migration
- Nach schwierigem Troubleshooting
### Format
`~/docs/solutions/{type}/{YYYY-MM-DD}-{slug}.md`
Types: `architecture/`, `bugfix/`, `migration/`, `workflow/`
Immer: Hindsight-Index (`hindsight_retain`) nach Erstellung
## Related
- [[systems/hindsight]] — Hindsight Setup, API
- [[systems/laya]] — Laya Classifier
- Skill: `memory-sync` — Wiki Maintenance Protocol
- Skill: `software-development/compound-learning` — Phase 4.7 Wiki Update
+24
View File
@@ -0,0 +1,24 @@
# User Drift Report — 2026-09-28
**Recent window:** Last 14 days (1884 msgs)
**Baseline:** Previous 90 days (6185 msgs)
## Detected Drifts
- **message_length**: Messages 98% longer (baseline: 1192 chars → recent: 2361)
- **new_focus**: New dominant topics: security
- **declining_focus**: Topics fading from focus: gitops
## Signal Summary
| Metric | Baseline | Recent |
|---|---|---|
| Avg msg length | 1192 | 2361 |
| Frustration rate | 4.2% | 6.6% |
| Correction rate | 7.1% | 8.8% |
| Action-first | 1.5% | 0.5% |
| Top topics | ceph, proxmox, k8s | security, backup, ceph |
## Recommendations
- Consider adding domain knowledge for new focus areas
+7 -4
View File
@@ -3,24 +3,27 @@ title: Infrastruktur-Übersicht
category: entities category: entities
tags: [homelab, hardware, overview] tags: [homelab, hardware, overview]
created: "2026-04-28" created: "2026-04-28"
modified: "2026-07-24" modified: "2026-09-26"
--- ---
# Infrastruktur-Übersicht # Infrastruktur-Übersicht
## Physikalische Hardware ## Physikalische Hardware
### Proxmox Cluster (PVE 9.2.3) ### Proxmox Cluster (PVE 9.2.20)
- **9 Nodes**, Quorum OK - **8 Nodes**, Quorum OK (proxmox1+2 dauerhaft entfernt)
- Hypervisoren in 10.0.20.x - Hypervisoren in 10.0.20.x
- Siehe [[systems/proxmox-cluster]] - Siehe [[systems/proxmox-cluster]]
### Ceph Cluster ### Ceph Cluster
- **14 OSDs** (HDD + NVMe/SSD混合) - **17 OSDs** (10 HDD + 7 SSD), all up/in, v20.2.4
- HEALTH_WARN (insecure key types — cosmetic)
- 4 MONs (px5 leader, px4, a2-1, n5pro), MGR auf n5pro
- Siehe [[systems/ceph-cluster]] - Siehe [[systems/ceph-cluster]]
### RKE2 Kubernetes Cluster ### RKE2 Kubernetes Cluster
- **6 Nodes** (3 CP + 3 Worker), v1.35.6-rke2r1 - **6 Nodes** (3 CP + 3 Worker), v1.35.6-rke2r1
- worker-04 currently NotReady (VM 139 offline)
- Alle Nodes schedulable (keine CP Taints) - Alle Nodes schedulable (keine CP Taints)
- Siehe [[systems/rke2-kubernetes]] - Siehe [[systems/rke2-kubernetes]]
+28 -2
View File
@@ -12,6 +12,7 @@
| `systems/` | Software-Systeme und Services | | `systems/` | Software-Systeme und Services |
| `concepts/` | Abstrakte Patterns & Konventionen | | `concepts/` | Abstrakte Patterns & Konventionen |
| `reference/` | Quick-Lookup Tabellen (IPs, Keys, Ports) | | `reference/` | Quick-Lookup Tabellen (IPs, Keys, Ports) |
| `patterns/` | Rekurrente Failure-Mode Patterns + Skill-Impact Tracker |
## Entities ## Entities
- [[entities/infrastructure]] — Physikalische Hardware, Cluster-Übersicht - [[entities/infrastructure]] — Physikalische Hardware, Cluster-Übersicht
@@ -21,8 +22,8 @@
- [[entities/health-fitness]] — Gesundheits-Ziele, Ernährung - [[entities/health-fitness]] — Gesundheits-Ziele, Ernährung
## Systems ## Systems
- [[systems/proxmox-cluster]] — PVE 9.2.3, 9 Nodes, Fluent Bit - [[systems/proxmox-cluster]] — PVE 9.2.20, 8 Nodes, Fluent Bit
- [[systems/ceph-cluster]] — 14 OSDs, EC Pools, bekannte Probleme - [[systems/ceph-cluster]] — 17 OSDs, v20.2.4, 4 MONs, bekannte Probleme
- [[systems/rke2-kubernetes]] — 6 Nodes, ArgoCD, CNPG, Workloads - [[systems/rke2-kubernetes]] — 6 Nodes, ArgoCD, CNPG, Workloads
- [[systems/galera-maxscale]] — 3 Galera + 2 MaxScale, VIP .70 - [[systems/galera-maxscale]] — 3 Galera + 2 MaxScale, VIP .70
- [[systems/loki-fluentbit]] — Logging Stack, 40+ Targets - [[systems/loki-fluentbit]] — Logging Stack, 40+ Targets
@@ -30,18 +31,43 @@
- [[systems/hindsight]] — Semantic Memory, K8s, API - [[systems/hindsight]] — Semantic Memory, K8s, API
- [[systems/monitoring]] — Prometheus, Grafana, HolmesGPT - [[systems/monitoring]] — Prometheus, Grafana, HolmesGPT
- [[systems/seafile]] — cloud.familie-schoen.com, Seafile 13.0.19 - [[systems/seafile]] — cloud.familie-schoen.com, Seafile 13.0.19
- [[systems/laya]] — 421M Classifier CT152, choice/noul/scale, Session-Pre-Filter
- [[systems/noris-ai]] — ai.noris.de LLMs, Embeddings, Image Gen
- [[systems/frigate]] — NVR v0.18 CT151, Kameras, MQTT, HA Automation
- [[systems/homeassistant]] — HAOS, MQTT, Automations, Notify
- [[systems/paperless]] — Dokumentenmgmt, OIDC, OCR via noris AI
- [[systems/sarah-hermes]] — VM107, zweites Hermes für Sarah (WebUI :8787, TG geplant)
## Concepts ## Concepts
- [[concepts/network-architecture]] — 10.0.X.Y Schema, VLANs - [[concepts/network-architecture]] — 10.0.X.Y Schema, VLANs
- [[concepts/gitops-workflow]] — IaC → Git → ArgoCD → Verify - [[concepts/gitops-workflow]] — IaC → Git → ArgoCD → Verify
- [[concepts/credential-policy]] — 1Password, ESO, keine Secrets in Git - [[concepts/credential-policy]] — 1Password, ESO, keine Secrets in Git
- [[concepts/email-organization]] — Rechnungs-Organizer, Himalaya - [[concepts/email-organization]] — Rechnungs-Organizer, Himalaya
- [[concepts/memory-layer-architecture]] — 7-Schichten Memory Trennungsregeln
## Reference ## Reference
- [[reference/ip-map]] — IP → Host → Service Mapping - [[reference/ip-map]] — IP → Host → Service Mapping
- [[reference/ssh-keys]] — Key → Zweck → Fingerprint - [[reference/ssh-keys]] — Key → Zweck → Fingerprint
- [[reference/ports]] — Port → Service → Host - [[reference/ports]] — Port → Service → Host
## Patterns
- [[patterns/_README]] — Konzept und Nutzung des Patterns-Verzeichnisses
- [[patterns/galera-ddl-deadlock]] — OPTIMIZE TABLE → TOI Deadlock (PAT-001)
- [[patterns/traefik-reload-unreliable]] — reload unreliable, restart statt reload (PAT-002)
- [[patterns/1password-rate-limit]] — Account-Level Rate-Limit, ESO stoppen (PAT-003)
- [[patterns/vfio-gpu-passthrough]] — softdep + DRM blacklist Race Condition (PAT-004)
- [[patterns/finanzblick-sync-waf]] — POST /sync WAF-blocked, UI Modal Sequence (PAT-005)
- [[patterns/linkedin-react-input]] — nativeInputSetter, Session Expiry (PAT-006)
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
- [[patterns/pve-oom-frozen-guest]] — Host-OOM friert Gast als Zombie ein, RSS-Kollaps-Signatur (PAT-010)
- [[patterns/pve-notification-webhook]] — PVE Webhook→Telegram: Endpoint-Drift, Base64-Trap, Handlebars-Escape (PAT-011)
- [[patterns/placement-guardrails]] — 70% RAM-Deckel + Swap-Verbot als Prometheus-Rules (PAT-012)
- [[patterns/prometheus-config-rescue]] — grep -vn Gift/Crashloop, LXC-Rootfs-Direct-Mount-Rescue (PAT-013)
- [[patterns/ceph-dead-nvme-resurrection]] — PCI-Reset→lvchange→chown-Falle, Weight-Management bei Reintegration (PAT-014)
- [[patterns/skill-impact]] — Skill Modification Audit Trail
## Memory Layer Architektur ## Memory Layer Architektur
| Layer | System | Pfad | Rolle | | Layer | System | Pfad | Rolle |
|-------|--------|------|-------| |-------|--------|------|-------|
+178 -2
View File
@@ -1,6 +1,174 @@
# Memory Log # Memory Log
## [2026-07-26] fix | Gitea SSH Push Key + Paperless IngressRoute Hostname ## [2026-09-30] k8s-cp01-relief-descheduler | Workload-Migration + Descheduler-Deployment
- cp-01 bei 93% RAM → hindsight-api + hindsight-postgres + paperless nach worker-05 migriert (cordon/delete/uncordon). cp-01 RAM 93%→43%.
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren bis konvergiert.
- hindsight-api Live-Pin (Sep-17-Hotfix) war nie in Git → Live-Patch nötig.
- Descheduler v0.30.0 deployt (raw Manifests, nicht Helm — Chart hat params/args-Bug + falsche API-Version). Policy: LowNodeUtilization (30/50/20 → 70/70/60), 2m Intervall, nodeFit auf Profile-Ebene.
- RBAC-Lektion: Descheduler braucht `list/watch` auf `namespaces` in ClusterRole.
- ArgoCD repoURL: `ssh://git@10.0.30.200:22/...` (mit `git@` User-Prefix, wie alle anderen Apps).
- Solution Doc + INDEX + Wiki rke2-kubernetes.md aktualisiert. Commits b955d7e→8c1ecff.
## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur
- Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen.
- Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate.
- SMART-Korrektur: Kingston SFYRDK2000G (osd.2-Träger) ist GESUND (wear 6 %, media_errors 0, spare 100 %, PoH 1469) — frühere "Austausch-Kandidatin"-These RETRAKIERT. Vorfall war Controller-/Fabric-Ebene. Neu beobachten: Temps 70/78 °C, T1-throttle 3×.
- PAT-014 erweitert (Smart-Log-Abschnitt), ceph-cluster.md korrigiert, Commit cc2f237.
## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014)
- osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin. *(Korrektur später am selben Tag: SMART-Audit → gesund, kein Austausch; siehe nächster Eintrag.)*
- osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend.
- ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben.
- Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014.
## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013)
- Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün.
- Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok.
- **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate.
- Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit <SHA>.
## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012)
- Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok.
- Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom).
- Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172.
## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011)
- Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON).
- Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200.
- Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body.
- Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales).
## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen
- proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%.
- Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets.
- VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP.
- Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt.
## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet
- GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md.
- Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc.
- proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen.
- Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert.
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX
- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared
- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr
- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return
- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010)
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
- E2E verifiziert: Login 200, Agent-Antwort via vllm/release/glm-5-2 @ ai.noris.de
- PITFALL: hermes-agent pyproject pinnt Core-Deps hinter python_version>='3.14'-Markers → auf Python 3.12 installiert `pip install -e .` NULL Deps ("AIAgent not available"). Fix: manuelle Pin-Installation + Missing-Import-Loop bis `import run_agent`
- Backup: täglich 23:00 Job backup-2e8a34e3-66cb (all=1) deckt VM107; manueller Verify TASK OK 48s
- Wiki: systems/sarah-hermes.md, ip-map.md erweitert; iac-homelab commits 0d5f017/1ce4816
- OFFEN: Telegram-Bot blockiert auf BotFather-Token (Sarah/Dominik)
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
- Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min)
- Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05
- Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2)
- Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes)
- Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
- Gitea Pod + schoenkitchen runner durch Pod-Delete wiederhergestellt
- ArgoCD Audit: 17/25 Apps via Gitea SSH, 20/25 Synced+Healthy
- Known Issues: `gitea`/`authelia` Health=Progressing (Ingress-LB-IP Gap), `immich` doppelt verwaltet, `gitea-config` Deployment Drift
- `DEFAULT_ACTIONS_URL=https://gitea.com` deprecated → Helm override auf `self` empfohlen
- Wiki `systems/gitea.md` aktualisiert: Runner Status, Incident, ArgoCD Issues
- Solution Doc: `docs/solutions/bug-fixes/2026-09-27-gitea-csi-rbac-volumeattachment-fix.md`
## [2026-09-27] architecture | Memory Layer Restructuring + Laya Session-Nutzung
- USER.md bereinigt: Infra-Fakten entfernt, nur noch User-Preferences/Safety-Rules (1.087/1.375 chars)
- MEMORY.md ausgedünnt: 2.145→1.664 chars, alle Infra-Details zeigen auf Wiki-Seiten mit `→Wiki` Pointern
- 4 neue Wiki-Seiten: systems/noris-ai, systems/frigate, systems/homeassistant, systems/paperless
- Neues Concept: concepts/memory-layer-architecture — Entscheidungsbaum "was wohin gehört"
- Hindsight Audit: massiv überladen mit veralteter Infra-Topologie. Going-Forward-Policy: nur noch semantische Pointer + Entscheidungen
- LCM gesund: 7.590 messages, 34 DAG nodes, 21.2:1 compression ratio
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
## [2026-09-27] architecture | Laya Email-Organizer Migration + Session-Nutzung
- Rechnungen-Organizer Cron `f773f8c23230` von LLM-Agent → `no_agent` Script mit Laya migriert
- 12 Kategorien, 3-Schichten-Safety (Confidence-Gate + Subject-Validierung + Move-Erfolg)
- Globale Ordner (Rechnungen/Bestellungen/Gutschriften/Gutscheine) statt Monatssortierung
- Wiki-Seite `systems/laya.md` erstellt mit Usage Guide für Session-Nutzung
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
## [2026-09-17] workflow | CT108-Endausbau: Census, Ghost-Router CT99999, Alias-Repair, Stop
- **Census (auth):** K8s=25 / CT108=31 / gemeinsam=23 — 22 Tips identisch, dominik/memory=K8s-Superset (enthält CT-Tip de97cd6d), nur-CT=8× PoC-Müll. Archiv: 29 Bare-Bundles 143 MB unter `/home/debian/git-archive/ct108-final/` (2 Failures = legitim leere Repos).
- **Ghost-Router enttarnt:** CT99999 (Traefik-LXC auf proxmox7, 10.0.60.10) terminiert TLS für *.familie-schoen.com und forwardet plain HTTP an .203. `git.familie-schoen.com` → .105:3000 WAR der letzte Live-Konsument von CT108; `git.schoen.codes` lief längst per Double-Hop aufs K8s. Fix: gitea-service-Upstream → .203 (Backup `explicit-http.yml.bak-hermes-20260917`) + Alias-Host im K8s-Ingress (Commit `9e9b6ee`). Beide Hostnamen jetzt v1.27.0.
- **Stop vollzogen (genehmigt):** `pct stop 108` auf ms-a2-2 (10.0.20.93) — connection-refused-Beweis, Fleet 22/25 Synced, Runner unversehrt. Wiki-Lügen korrigiert (ip-map behauptete „stopped" seit Wochen).
- **Lessons:** 1P-SA braucht je Call `--vault`; `op read --reveal` existiert nicht (stdout=Secret); `op item get --reveal` maskiert nur Display (JSON-Captures intakt); `/repos/search` = `{ok,data}`-Envelope; blankes Token erzeugt glaubwürdig LEERE Census (Fast-Fehlentscheidung „K8s hat nur 2 Repos").
- Docs: `docs/solutions/workflows/2026-09-17-ct108-full-decommission-census-router-topology.md` · Wiki: systems/gitea.md, reference/ip-map.md
- Offen (je Freigabe): Zombie-Secret `argocd-repo-credentials` löschen; immich-Zwillings-App (toter rendered-Dump vom 01.08., Live gehört immich-config) entfernen.
## [2026-09-17] fix | Merge-Day-Kampagne abgeschlossen: Konvergenz + Autosync-Rennen + Live-State-Arbitrage
- **Konvergenz DONE:** Merge `d76d3e8` (fork 46fa168, 24.07.) + Fixups `4987341`/`8c741c5` auf BEIDEN Remotes (CT108 + K8s-Gitea SSH 10.0.30.200). Ahead37/behind63-Narrativ endgültig begraben (Cache-Phantom). Single-Remote-Ziel erreicht: beide Tipps identisch.
- **Autosync-Renne verarbeitet (4 Minentypen):** authelia Duplikat-Volume (union-merge) entfernt; homepage `authelia-auth`-Middleware restauriert (Jul-Entscheidung ging nie live); paperless plaintext-OIDC-Secret GELÖSCHT statt Wert-Rollback (ESO-Ownership seit 24.08., Quelle 1P); newborn-App ceph-csi-cephfs eingefroren (autosync-Block aus Git entfernt — Helm-Release 3.17.0 wartet auf Adoption; rbd-chart 3.10.1 weiterhin ungoverned, Backlog).
- **3 Sync-Failures via Live-State-Arbitrage geheilt (alle Synced/Healthy @ 8c741c5):** (1) backups: VSC-CRD-Flavor-Falle — RKE2-Addon-CRD ist FLAT (driver/deletionPolicy top-level, `spec:` verboten!), Jul-Files waren für Upstream-Flavor korrekt; beide Files geflattet, cephfs-snapclass erstmals LIVE ERZEUGT. (2) databases: CNPG grow-only — Git auf 100Gi/20Gi hochaligned (Shrink verboten). (3) hindsight: VCT storageClassName immutable (hdd→flash aligniert), Live-Hotfixes gespiegelt (Node-Pin cp-01, ReadinessProbe draußen, secretKeyRef statt Klartext-PW). hindsight-postgres-0 rotierte sauber, PVC 47d Bound, API pollt.
- **Push-Learned:** canonical-Push braucht EXPLIZITEN Key (`GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o IdentitiesOnly=yes"`) — kein ~/.ssh/config vorhanden, Default-Key lehnt ab.
- **Chronisch (prä-merge, offen):** kube-prometheus-stack Synced/FAILED (CRD-Annotation >256kB); residual OutOfSync gitea-config (Deployment/gitea-runner) + immich-config (SA/CM-Reste) = Alt-Backlog, keine Regression.
- **Finale Ordnung (nächste Schritte):** ① ArgoCD-Sources auf ssh://git@10.0.30.200:22 flippen (incl. Secret argocd-repo-credentials) ② CT108 final stoppen (explizite Freigabe Dominik) ③ CSI-Adoption ④ kpstack-CRD-Fix.
- Docs: `docs/solutions/workflows/2026-09-17-argocd-autosync-race-four-mine-types.md` + `docs/solutions/bug-fixes/2026-09-17-live-state-arbitration-crd-flavors-hotfix-mirroring.md` · Morgen-Doc (.202-Empfehlung) korrigiert auf .200.
## [2026-09-17] bug-fix | Ghost-Instance-Regression: CT108-Gitea als ArgoCD-Source + gitea-backup RCA-Quality
- gitea-backup-29826930 (03:30Z) failed: BackoffLimitExceeded, 3 Instant-Crashes in 73s. Zwei Auto-RCAs attribuierten auf 1P/ESO-Rate-Limit — STRUKTURELL widerlegt (crashing init-container gitea-files konsumiert keine Secrets; mounted Secrets intakt).
- Echter Fund: ArgoCD-Apps (root/proxy/paperless/schoenkitchen/gitea-config) zogen von http://10.0.30.105:3000 = CT108-Ghost (nach Dekommissionierung 24.07. wieder eingeschaltet, stale Mirror ohne Hardening-Commit 2014673) → "Synced" maskierte fehlende failedJobsHistoryLimit-Felder live.
- Sofortmassnahme: manueller Rerun gitea-backup-manual-161538 SUCCESS (201,7 MB, S3-Upload verifiziert), Failed-Job gelöscht, Alert clear.
- Offen (awaiting owner): ArgoCD-Sources auf ssh://git@10.0.30.202:22 flippen, Repo-Divergenz ahead37/behind63 + tote HTTPS-Auth forensisch, CT108 final stoppen, ESO-Nachtsättigung (00:00Z-Fenster, cf. PAT-003) analysieren.
- Doc: `docs/solutions/bug-fixes/2026-09-17-gitops-ghost-instance-regression-masked-hardening.md` · Wiki: systems/gitea (CT108-Status), concepts/gitops-workflow (Single-Remote-Regel)
- **Abend-Phase:** SSH-Deny-Rootcause = keine Keys/Tokens im K8s-Gitea-DB registriert (Opfer der 01.08.-Migration) → via 1P-Token (hermes-gitops) re-registriert: User-Key + ArgoCD-Deploy-Key (beide end-to-end verifiziert, push dry-run OK). Echter Branch-Split am 24.07. entdeckt: K8s-Gitea-main=29.07.-Stand (b9d7441, 70 Commits incl. mariadb:11.4-Fix für gitea-backup), CT108/local=38 Commits ab 04.09. Hardening TTL=86400 + Exit-42-Guard via ArgoCD gelanded (a0c9b6b). Heutiger DB-Dump validiert (116 Tables, kompletter Trailer). Nächste Schritte: Historien-Konvergenz → Source-Flip → CT108-Stop (Freigabe).
## [2026-08-30] retro | Compound Learning Retrospective (Last 30 Days)
- Reviewed sessions from Jul 31 – Aug 30, 2026
- **4 new solution docs** written by subagents:
- `bug-fixes/2026-08-05-ceph-squid-ec-pool-mark-complete-bug.md` — EC pool mark-complete fails in Squid
- `bug-fixes/2026-08-06-traefik-cross-namespace-routing-pitfall.md` — Two Traefik 404 incidents
- `architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md` — Full RKE2 rebuild sequence
- `architecture/2026-08-01-immich-data-loss-no-backup.md` — Ceph redundancy ≠ backup
- **2 new patterns** extracted:
- PAT-008: Ansible default_ipv6 fact missing on fresh VMs
- PAT-009: Stale NBD devices after RBD volume swap
- **4 Hindsight entries** indexed with solution summaries
- Key themes: K8s cluster disaster recovery, Ceph EC pool unrecoverability, Traefik routing complexity, backup gap identification
## [2026-08-30] feat | WikiSkill Patterns Directory + Skill-Impact Tracker
- Analysed arXiv:2608.27454 (WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution)
- **New `patterns/` directory** in LLM Wiki — structured failure-mode patterns inspired by WikiSkill's Wiki Layer
- Created 7 initial pattern pages from existing MEMORY.md entries + solution docs:
- PAT-001: Galera DDL TOI Deadlock
- PAT-002: Traefik reload unreliable
- PAT-003: 1Password account-level rate-limit
- PAT-004: VFIO GPU passthrough race condition
- PAT-005: Finanzblick sync WAF-blocked
- PAT-006: LinkedIn React nativeInputSetter
- PAT-007: Ceph EC pool + SSD wear-level
- **Skill-Impact Tracker** (`patterns/skill-impact.md`) — audit trail for skill modifications (accept/reject history)
- **Template** (`patterns/_template.md`) for future pattern creation
- **compound-learning skill patched** — added Phase 4.8 (Pattern Extraction) + Phase 4.9 (Skill-Impact Tracking)
- Updated index.md with Patterns section
- Hindsight: indexed with tags architecture, wikiskill, patterns, skill-evolution
## [2026-08-13] feat | Memory Architecture Enhancement (arxiv 2602.06052v4)
- Implemented 3 new automated memory capabilities from survey paper analysis:
- **Decay & Importance Scoring** (`hindsight_decay_scoring.py`): Monthly cron (1st 03:30). Score = log(1+access)/log(51) × 0.5^(age/90d). Archives score<0.15 with 0 accesses. First run: 200 evaluated, 50 archived.
- **User Drift Detection** (`user_drift_detection.py`): Weekly cron (Mon 04:00). Compares 14d vs 90d baseline for message length, frustration, corrections, topic shifts, style changes. First run: detected 3 drifts (41% longer msgs, new focus gitops, declining ai_ml).
- **Skill Health Check** (`skill_health_check.py`): Weekly cron (Mon 04:30). Scans 187 active SKILL.md files for stubs, broken refs, missing pitfalls, staleness, duplicates. First run: 70 healthy, 117 warnings, 0 critical.
- Cron jobs: 112c106c0954, 280458a2b035, 761b6694d99b (all no_agent, deliver=origin)
- Wiki updated: systems/hindsight.md (meta-memory automation table expanded)
- Hindsight: indexed with tags architecture, memory-system, cron, arxiv-2602.06052
## [2026-08-13] fix | Seafile Recovery + Paperless Upgrade + Authelia Ingress
- **Gitea SSH access**: No SSH key on Hermes VM was authorized for Gitea. HTTP port 3000 has no external IngressRoute. Generated dedicated ED25519 key (`id_ed25519_gitea-hermes-push`), stored in 1Password (ID: lqaebdtewj5xatvpzfrkag73v4), registered in Gitea as key ID 4 for user dominik. SSH via `10.0.30.202:22` (LoadBalancer) is the reliable push path. - **Gitea SSH access**: No SSH key on Hermes VM was authorized for Gitea. HTTP port 3000 has no external IngressRoute. Generated dedicated ED25519 key (`id_ed25519_gitea-hermes-push`), stored in 1Password (ID: lqaebdtewj5xatvpzfrkag73v4), registered in Gitea as key ID 4 for user dominik. SSH via `10.0.30.202:22` (LoadBalancer) is the reliable push path.
- **Paperless v3 IngressRoute**: Host was `dokumente.familie-schoen.com` instead of `dokumente-neu.familie-schoen.com`. Two-track fix: kubectl patch (immediate) + Git commit `fa4184b` (permanent via ArgoCD self-heal). Backtick escaping in Traefik Host() match requires `--patch-file` not inline `--patch`. - **Paperless v3 IngressRoute**: Host was `dokumente.familie-schoen.com` instead of `dokumente-neu.familie-schoen.com`. Two-track fix: kubectl patch (immediate) + Git commit `fa4184b` (permanent via ArgoCD self-heal). Backtick escaping in Traefik Host() match requires `--patch-file` not inline `--patch`.
- Wiki updated: systems/gitea.md (SSH push key section), reference/ssh-keys.md (new key entry) - Wiki updated: systems/gitea.md (SSH push key section), reference/ssh-keys.md (new key entry)
@@ -27,7 +195,7 @@
- Velero backups PartiallyFailed for 121 days — all PVCs skipped - Velero backups PartiallyFailed for 121 days — all PVCs skipped
- Fix 1: Created VolumeSnapshotClass `ceph-rbd-snapclass` (rbd.csi.ceph.com) - Fix 1: Created VolumeSnapshotClass `ceph-rbd-snapclass` (rbd.csi.ceph.com)
- Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22) - Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22)
- Gitea SSH user is `git` not `gitea`; git.schoen.codes → Traefik (10.0.30.203), SSH on 10.0.30.202 - Gitea SSH user is `git` not `gitea`; git.schoen.codes → Traefik (10.0.30.208), SSH on 10.0.30.202
- New SSH deploy key generated, added to Gitea repo - New SSH deploy key generated, added to Gitea repo
- Fix 3: `features: EnableCSI` must be under `configuration:` in Velero Helm values (not top-level) - Fix 3: `features: EnableCSI` must be under `configuration:` in Velero Helm values (not top-level)
- Must be string, not array — array form breaks Helm template - Must be string, not array — array form breaks Helm template
@@ -122,3 +290,11 @@
- Created 2 concept pages: gitops-workflow, credential-policy - Created 2 concept pages: gitops-workflow, credential-policy
- Rewrote index.md with new structure + Memory Layer Architecture table - Rewrote index.md with new structure + Memory Layer Architecture table
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages) - Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
## [2026-09-29] update | VM302 Zombie-Recovery nach Migration
- VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot
- Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause
- Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min
- Endstand: 3/3 Synced, Primary, MaxScale alle Server Running
- Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md
- Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)
+51
View File
@@ -0,0 +1,51 @@
---
pattern_id: PAT-003
title: "1Password CLI rate-limit is account-level — stop ESO, wait 1 hour"
category: tooling
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [rke2-kubernetes]
related_solution_docs:
- docs/solutions/architecture/2026-07-24-1password-account-level-rate-limit.md
- docs/solutions/bug-fixes/2026-07-14-1password-token-telegram-corruption.md
related_skills: [1password-cli]
---
# 1Password CLI rate-limit is account-level — stop ESO, wait 1 hour
## Symptom
All 1Password CLI (`op`) calls suddenly fail with rate-limit errors. This affects
both Hermes tools and ExternalSecrets Operator (ESO) in K8s simultaneously.
The outage lasts approximately 1 hour.
## Root Cause
The 1Password CLI enforces rate limits at the **account level**, not per-token or
per-client. When ESO polls aggressively (multiple SecretStores + frequent sync intervals),
it exhausts the quota for the ENTIRE account, blocking all other `op` callers including
Hermes automation.
## Mitigation
1. **Stop ESO immediately** — scale down the ESO deployment to halt API calls:
```bash
kubectl scale deploy -n external-secrets external-secrets-operator --replicas=0
```
2. **Wait 1 hour** — the rate limit resets automatically. No polling needed.
3. **Resume ESO** with reduced polling frequency afterward
## Prevention
- Reduce ESO `refreshInterval` to 10+ minutes (not the default 1 minute)
- Limit the number of SecretStores that reference 1Password
- Consider caching secrets locally to reduce API pressure
- Never run `op` in tight loops — always add delays for batch operations
## Evidence
- Hit twice: Jul 2026 (initial discovery) and during token corruption investigation
- Solution doc: `docs/solutions/architecture/2026-07-24-1password-account-level-rate-limit.md`
- MEMORY.md entry: "1P CLI:3 vaults R/W. Rate-limit=acct-level. Stop ESO, wait 1hr, no polling."
+51
View File
@@ -0,0 +1,51 @@
# Patterns Directory
Inspired by [WikiSkill (arXiv:2608.27454)](https://arxiv.org/abs/2608.27454) — this directory
contains structured **failure-mode patterns** and **successful strategies** extracted from
real operational experience.
## Purpose
Unlike `systems/` (describes what exists) or `concepts/` (abstract conventions),
`patterns/` captures **recurrent problems and their proven solutions** — the "lessons learned"
that compound across incidents.
## Structure
Each pattern is a standalone Markdown file:
```
patterns/
├── _README.md ← this file
├── _template.md ← copy this for new patterns
├── galera-ddl-deadlock.md
├── traefik-reload-unreliable.md
├── ...
└── skill-impact.md ← audit trail of skill modifications (accept/reject history)
```
## How Patterns Are Born
1. **Incident occurs** → problem is diagnosed and fixed
2. **Compound-learning cycle** runs → solution doc created in `~/docs/solutions/`
3. **Pattern extracted** → if the failure mode is recurrent or broadly applicable,
a pattern page is created here
4. **Wiki Maintainer** (currently manual, future: automated) consolidates evidence
from multiple incidents into the pattern page
## Relationship to Other Layers
| Layer | Holds | Retrieval |
|-------|-------|-----------|
| `~/docs/solutions/` | One-time detailed incident docs | `search_files` (keyword) |
| `patterns/` (here) | Recurring failure-mode patterns | Browse `index.md`, cross-linked |
| MEMORY.md (L0) | Compressed pointer if high-priority | System prompt injection |
| Hindsight (L2) | Semantic index of all above | `hindsight_recall` |
## Rules
- **One pattern per file** — don't merge unrelated patterns
- **Evidence-based** — cite real incidents (link to solution docs or session dates)
- **Actionable** — every pattern must have a "Mitigation" or "Prevention" section
- **Never delete without redirect** — if a pattern is obsolete, mark `status: superseded`
and link to the replacement
+30
View File
@@ -0,0 +1,30 @@
---
pattern_id: PAT-XXX
title: "<short descriptive title>"
category: database|infrastructure|storage|networking|tooling|integration|security
severity: low|medium|high
status: active|superseded
first_observed: YYYY-MM
last_updated: YYYY-MM-DD
related_systems: []
related_solution_docs: []
related_skills: []
---
# <Title>
## Symptom
What goes wrong? What are the observable symptoms?
## Root Cause
Why does it happen? Trace the actual cause, not just the symptom.
## Mitigation
What was the fix? Include commands/snippets if relevant.
## Prevention
How to avoid this in the future? (monitoring, lint rule, convention, etc.)
## Evidence
- When was this observed? Which incidents?
- Links to solution docs, session IDs, MEMORY entries
+59
View File
@@ -0,0 +1,59 @@
---
pattern_id: PAT-008
title: "Ansible default_ipv6 fact missing on fresh VMs → lablabs.rke2 role fails"
category: tooling
severity: medium
status: active
first_observed: 2026-07-21
last_updated: 2026-08-30
related_systems: [rke2-kubernetes]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-21-ansible-rke2-default-ipv6-bug.md
related_skills: []
---
# Ansible default_ipv6 fact missing on fresh VMs → lablabs.rke2 role fails
## Symptom
The `lablabs.rke2` Ansible role fails on freshly provisioned VMs with a Jinja2 error:
```
ansible_facts['default_ipv6']['address']
```
The `default_ipv6` fact doesn't exist on VMs without IPv6 configured, causing the
`meta/argument_specs.yml` validation to crash.
## Root Cause
The `lablabs.rke2` role's `meta/argument_specs.yml` references
`ansible_facts['default_ipv6']['address']` with a hardcoded default. On fresh VMs
without IPv6, `ansible_facts['default_ipv6']` is undefined → Jinja2 raises
`UndefinedError`.
A `pre_task` setting `default_ipv6` via `combine()` doesn't help because the
argument_specs validation runs BEFORE pre_tasks — it re-collects facts and
overwrites the pre_task fix.
## Mitigation
Patch `meta/argument_specs.yml` directly in the role:
```yaml
# Replace:
default: "{{ ansible_facts['default_ipv6']['address'] }}"
# With:
default: "{{ ansible_facts.default_ipv6.address | default(None) }}"
```
Or: enable IPv6 on the target VMs (faster workaround, no role patching needed).
## Prevention
- Pin Ansible roles and patch argument_specs when they assume facts that may not exist
- Test roles on fresh VMs without IPv6 before production use
- Consider forking the role with the fix upstream
## Evidence
- Observed during GPU worker provisioning (worker-04/05, Jul 2026)
- Solution doc: `docs/solutions/bug-fixes/2026-07-21-ansible-rke2-default-ipv6-bug.md`
- Session: @session:default/20260721_115501_78b25032
+44
View File
@@ -0,0 +1,44 @@
# PAT-014 — Dead-NVMe-Resurrection für Ceph-OSD
## Trigger
Ceph-OSD auf externem/non-Corosync-Host startet nicht: KernelDevice errno-5 I/O-Errors beim
Label-Lesen, Symlink `/var/lib/ceph/osd/ceph-*/block` ins Leere, VG „verschwindet".
## Root Cause
NVMe-Controller im PCIe-Fabric gestorben (state=dead, Namespace 0B) — NICHT das Medium selbst.
Nach PCI-reset kehrt der Controller oft unter NEUER Nummer zurück (nvme0 → nvme2!), alte
Device-Mapper-Tabelle bleibt stale.
## Resurrection-Sequence (verifizierte Reihenfolge)
1. **Korrekte PCI-Addr finden**: über sysfs-Pfad des Namespaces (`/sys/class/block/nvmeXnY/device`),
NICHT raten — erster Versuch traf den falschen (gesunden!) Controller.
2. `echo 1 > /sys/bus/pci/devices/<ADDR>/remove && echo 1 > /sys/bus/pci/rescan`
→ Controller kommt als neue Instanz zurück, Namespaces wieder da.
3. `pvscan --cache` → VG/LV wieder sichtbar.
4. **Stale DM-Table**: `lvchange -an <vg>/<lv> && lvchange -ay -K <vg>/<lv>`
5. **Lesetest**: `dd if=/dev/<vg>/<lv> bs=4M count=8 of=/dev/null` — bei weiteren errno-5:
NAND/Media defekt → OSD out lassen, Drive tauschen.
6. **Permissions-Falle nach lvchange**: udev setzt Owner root →
`chown ceph:ceph /dev/mapper/<dm-name>; chmod 660` (sonst `bdev open: (13) Permission denied`).
7. `systemctl reset-failed ceph-osd@<id> && systemctl start ceph-osd@<id>`
8. `ceph osd in osd.<id>` + Weight restaurieren.
## Begleitregeln
- Klein/niedergewichtetes OSD niemals mit Weight 1.0 reintegrieren, solange Pools an
Full-Ratios kratzen → sonst `backfillfull`/`backfill_toofull`-Blockade (Fall osd.3:
re-in@1.0 → 96 % voll instant → 13 PGs blocked; Fix: `crush reweight osd.3 0.05`).
- Externe Non-Corosync-Hosts in `/etc/hosts` ALLER PVE-Nodes pflegen (GUI-500
„hostname lookup failed"), oder echter DNS-Record.
## Verified
2026-09-30: osd.2 revived (Kingston SFYRDK2000G, PCI 03:00.0, → nvme2), 495 GiB, Weight 1.0;
Cluster 17/17 up/in; Degraded 0,78 % fallend. osd.3 stabilized @ weight 0.05.
## Smart-Log-Abgleich 2026-09-30 (korrigiert frühere Alters-These)
Kingston SFYRDK2000G (nvme0n2, trägt osd.2): percentage_used **6 %**, media_errors **0**,
available_spare 100 %, unsafe_shutdowns 2, power_on_hours 1469, power_cycles 3.
Geschwisterplatte nvme1n1 identisches Profil (6 % wear, 0 media errors).
⇒ Drive ist MEDIZINISCH GESUND — der Vorfall war rein Controller-(PCI)-Ebene, kein
Media-Verschleiß. **Kein Austausch nötig.** Einziges Beobachtungsfeld: Temperatur 70 °C /
Sensor2 78 °C (thermisches Throttling T1 3× aktiviert) — Kühlung prüfen wäre sinnvoll,
aber keine Akutgefahr.
+60
View File
@@ -0,0 +1,60 @@
---
pattern_id: PAT-007
title: "Ceph EC pool unusable + SSD wear-level failing — monitor pg states"
category: storage
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [ceph-cluster]
related_solution_docs:
- docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md
- docs/solutions/bug-fixes/2026-07-13-ceph-ratio-ordering-constraint.md
- docs/solutions/bug-fixes/2026-07-13-pveceph-wal-lvm-device-mapper-limitation.md
related_skills: [ceph-cluster-administration, ceph-ec-incomplete-pg-recovery]
---
# Ceph EC pool unusable + SSD wear-level failing — monitor pg states
## Symptom
Multiple Ceph failure modes manifest simultaneously:
1. EC pool (pool 8/media_ec) becomes unusable — PGs stuck in `incomplete` state
2. SSD OSD (osd.3 on px4) reports `Wear Leveling Count` SMART attribute indicating
imminent failure
3. `ceph health` shows `HEALTH_ERR`
## Root Cause
These are compounded issues:
1. **EC pool incomplete PGs**: After OSD loss, EC pools cannot recover without enough
surviving shards. The minimum copies requirement for EC k+m encoding is stricter than
replicated pools.
2. **SSD wear-level failure**: Consumer-grade SSDs in Ceph OSD duty cycle reach wear limits.
Once `Wear Leveling Count` exceeds threshold, the SSD becomes read-only or unreliable.
3. **Weight imbalance**: Unequal OSD weights cause CRUSH placement failures, leaving PGs
perpetually remapped.
## Mitigation
1. **EC pool recovery**: Use the `ceph-ec-incomplete-pg-recovery` skill — specialized
procedure for forcing EC PG recovery after OSD loss
2. **Replace failing SSD**: Mark OSD `out` → `destroy` → `zap` → physically replace → recreate
3. **Fix weight imbalance**: Equalize `pg_num` → `pgp_num`, set `nopgchange=true`,
reweight OSDs proportionally to disk capacity
4. **Remove empty CRUSH buckets**: Clean up stale host entries from decommissioned nodes
## Prevention
- Monitor SMART attributes monthly — alert on wear-level > 80%
- Never use consumer SSDs for Ceph OSDs in production (use enterprise grade with PLP)
- Keep EC pool k+m ratio proportional to available OSDs (don't over-provision redundancy)
- Regular `ceph pg dump` audits for stuck/unactive PGs
- Separate PBS onto its own tier (CephFS), away from RBD pools
## Evidence
- Multiple incidents Jul 2026 (OSD 0+2 destruction, OSD 3 wear-level, EC pool unusable)
- Solution docs: `2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`,
`2026-07-13-ceph-ratio-ordering-constraint.md`
- MEMORY.md entry: "Ceph:HEALTH_ERR. Pool8/media_ec unusable. px4 SSD(osd.3) WearLevel FAILING."
+53
View File
@@ -0,0 +1,53 @@
---
pattern_id: PAT-005
title: "Finanzblick sync requires modal sequence — POST /sync is WAF-blocked"
category: integration
severity: medium
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: []
related_solution_docs: []
related_skills: [finanzblick-cashflow]
---
# Finanzblick sync requires modal sequence — POST /sync is WAF-blocked
## Symptom
Programmatic synchronization with Finanzblick (banking data aggregator) fails when
calling the `POST /sync` API endpoint directly. The request is blocked by the WAF
(Web Application Firewall), returning 403 or connection reset.
Historical fetches (without sync) work fine with the `--no-sync` flag, avoiding 2FA.
## Root Cause
Finanzblick's WAF detects and blocks automated POST requests to the sync endpoint
that don't originate from the legitimate browser session with proper CSRF tokens
and session cookies.
## Mitigation
Sync must be performed via the **UI button + 2FA modal sequence**:
1. Navigate to the Finanzblick web interface in a browser
2. Click the sync button (UI-triggered, not API)
3. Handle the 2FA modal sequence in order:
- PIN modal → click OK
- AUTH modal → click WEITER
- ERR modal → click OK
4. Wait for sync completion
For historical data fetches (no sync needed), use the `--no-sync` flag — this
bypasses 2FA entirely.
## Prevention
- Never attempt direct `POST /sync` calls — always use the UI flow
- The `finanzblick-cashflow` skill encodes this modal sequence
- This skill is USER-OWNED and needs `hermes curator adopt` to manage
## Evidence
- Observed during Finanzblick cashflow analysis sessions (Jul 2026)
- MEMORY.md entry: "FB sync=UI btn+2FA modals(PIN→OK,AUTH→WEITER,ERR→OK). POST /sync=WAF-blocked. --no-sync flag for hist.fetches(no 2FA). fb-cashflow skill=USER-OWNED,needs `hermes curator adopt`."
+52
View File
@@ -0,0 +1,52 @@
---
pattern_id: PAT-001
title: "OPTIMIZE TABLE on Galera triggers TOI DDL → Deadlock"
category: database
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [galera-maxscale]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-23-cnpg-timeline-corruption-rebuild.md
related_skills: []
---
# OPTIMIZE TABLE on Galera triggers TOI DDL → Deadlock
## Symptom
Running `OPTIMIZE TABLE` on a Galera cluster node causes a cluster-wide stall or deadlock.
The table lock propagates via TOI (Total Order Isolation) to all nodes, blocking all writes
to ALL tables during the operation.
## Root Cause
Galera executes DDL statements (including `OPTIMIZE TABLE`, `ALTER TABLE`, etc.) in TOI mode.
This means the DDL is replicated as a global operation that blocks the entire cluster — not
just the target table. For large tables, the rebuild phase can take minutes, causing apparent
outages.
## Mitigation
1. **Use RSU before DDL**: Switch the node to Rolling Schema Update (RSU) mode before running
`OPTIMIZE TABLE`. This prevents cluster-wide blocking:
```
SET GLOBAL wsrep_OSU_method = 'RSU';
-- run OPTIMIZE TABLE on this node only
SET GLOBAL wsrep_OSU_method = 'TOI'; -- restore
```
2. **Schedule during maintenance window** — even with RSU, the node itself is degraded
3. **Consider pt-online-schema-change** for large tables — avoids blocking entirely
## Prevention
- Never run `OPTIMIZE TABLE` or `ALTER TABLE` in production without RSU on Galera
- Add this check to DBA runbooks and monitoring alerts
- The Galera skill (`mariadb-galera-cluster-administration`) documents this pitfall
## Evidence
- Observed during Galera cluster administration sessions (Jul 2026)
- MEMORY.md entry: "OPTIMIZE TABLE=TOI DDL→Deadlock! RSU davor."
- Galera cluster: nodes 300/301/302, VIP .70:3306
+66
View File
@@ -0,0 +1,66 @@
---
pattern_id: PAT-009
title: "Stale NBD devices after RBD volume swap — delete VolumeAttachment + nbd teardown"
category: infrastructure
severity: high
status: active
first_observed: 2026-08-01
last_updated: 2026-08-30
related_systems: [rke2-kubernetes, ceph-cluster]
related_solution_docs:
- docs/solutions/architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md
related_skills: []
---
# Stale NBD devices after RBD volume swap — delete VolumeAttachment + nbd teardown
## Symptom
After swapping RBD images for PVCs (e.g. migrating to new Ceph pool or recreating
images), pods fail to mount with errors like:
```
MountVolume.MountDevice failed for volume "pvc-xxx" : rpc error: code = Internal
desc = rbd: map failed with error: /dev/nbd0 already in use
```
The NBD device is held by a stale mapping from the old RBD image, even though the
new image has the same name.
## Root Cause
When an RBD image is recreated (delete + create with same name), the Ceph CSI
driver's NBD mappings from the old image remain active. The Linux NBD layer
holds `/dev/nbdX` open, blocking new mounts to the same device path.
The Kubernetes VolumeAttachment object also references the old volume handle,
preventing the CSI driver from cleanly attaching the new volume.
## Mitigation
Three-step teardown procedure:
```bash
# 1. Delete the VolumeAttachment (allows CSI driver to release)
kubectl delete volumeattachment csi-cephfsplugin-<node>-<volume-handle>
# 2. Disconnect the stale NBD device on the target node
ssh <node> 'nbd-client -d /dev/nbd0' # or: qemu-nbd --disconnect /dev/nbd0
# 3. Restart the CSI node plugin to pick up clean state
kubectl delete pod -n kube-system csi-cephfsplugin-<node-id>
# (DaemonSet will respawn it)
```
After this, the pod can remount with the new RBD image.
## Prevention
- Before deleting RBD images, ensure all pods using them are scaled to 0
- Delete VolumeAttachments BEFORE deleting RBD images
- After RBD image recreation, restart CSI plugins on all nodes that had mounts
- Document this in the K8s disaster recovery runbook
## Evidence
- Observed during K8s cluster rebuild (Aug 2026) — 13 RBD volumes needed swapping
- Session: @session:default/20260731_113701_74fd2814 (250+ tool calls)
- Part of the full cluster rebuild procedure
+58
View File
@@ -0,0 +1,58 @@
---
pattern_id: PAT-006
title: "LinkedIn login via browser — nativeInputSetter required, session expires between nav"
category: integration
severity: medium
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: []
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-10-ha-token-expiry-browser-session-loss.md
related_skills: [linkedin-personal-branding]
---
# LinkedIn login via browser — nativeInputSetter required, session expires between nav
## Symptom
Automated LinkedIn login via browser tools fails silently. Form fields appear filled
but LinkedIn doesn't recognize the input (login button stays disabled, or submission
fails with "invalid credentials").
Additionally, authenticated sessions expire between page navigations, requiring
re-login on almost every navigation step.
## Root Cause
1. **React-controlled inputs**: LinkedIn uses React, which manages input state internally.
Simply setting `.value` via DOM manipulation doesn't trigger React's onChange handler.
The `nativeInputSetter` approach is required:
```javascript
const setter = Object.getOwnPropertyDescriptor(
window.HTMLInputElement.prototype, 'value'
).set;
setter.call(inputElement, 'my-value');
inputElement.dispatchEvent(new Event('input', { bubbles: true }));
```
2. **Session expiry**: LinkedIn's SPA doesn't persist auth tokens reliably across
`browser_navigate` calls (each navigation may reset the JS context).
## Mitigation
- Always use `browser_console` with `nativeInputSetter` for form fills — NOT `browser_click`
followed by typing
- Perform login + desired action in a SINGLE browsing session (minimize navigations)
- UTF-8 files must be saved without BOM (LinkedIn rejects BOM in posted content)
## Prevention
- The `linkedin-personal-branding` skill documents this workflow
- Never use `browser_click` for LinkedIn form fields
- Minimize navigation steps after login
## Evidence
- Observed during LinkedIn branding sessions (Jul 2026)
- MEMORY.md entry: "Login via browser_console nativeInputSetter. NOT browser_click. Session expires between nav. UTF-8 no BOM."
+23
View File
@@ -0,0 +1,23 @@
# PAT-012 — Placement Guardrails (RAM-Deckel + Swap-Verbot)
**Trigger**: Neue Gastplatzierung oder Kapazitätsfrage auf einem PVE-Node.
**Policy**:
- Committed RAM pro Hypervisor ≤ 70 % physikalisch. Über Deckel = "voll",
Guests redistribuieren BEVOR neue Last kommt.
- Resident Swap > 100 MiB für >10 min = Politikverstoß → Placements prüfen,
nie als Normalzustand akzeptieren.
**Enforcement (live)**: Gruppe `placement_guardrails` in CT141
(`/opt/monitoring/prometheus/rules/alerting_rules.yml`):
- `PVEPlacementCeilingBreached` (ratio > 0.70, for 15m, warning)
- `HVResidentSwapNonzero` (SwapTotal-Free > 100MiB, for 10m, warning)
**Vor jeder Platzierung**: Ratio-Query ziehen
(`pve_memory_usage_bytes{id=~"node/.*"} / pve_memory_size_bytes{id=~"node/.*"}`)
— >70 %-Nodes meiden, Headroom liegt primär auf ms-a2-1/ms-a2-2.
Baseline-Rollout 30.09.: proxmox3 80,6 / p4 79,3 / p5 75,5 / p6 72,6 %
über Deckel (Warnburst erwartet = Rebalancing-Backlog, kein Emergency).
Ref: docs/solutions/architecture/2026-09-30-placement-guardrails-ram-ceiling.md
+28
View File
@@ -0,0 +1,28 @@
# PAT-013 — Prometheus Config Rescue (Crashloop + LXC Rootfs Channel)
**Trigger**: Prometheus crashloop nach Config-Rewrite, oder pct exec unbrauchbar.
**Iron Rules**:
1. NIEMALS `grep -n`-Ausgabe als Rewrite-Source — Nummernpräfixe verseuchen die Datei.
Nutze `grep -v` (ohne -n) oder Python/awk.
2. YAML validieren BEVOR Container recreate (sonst Crashloop = totale Blindheit).
3. Beim Löschen von Targets/IPs ALLE Locations sweeppen — Dubletten leben gern in
mehreren Jobs derselben Datei (node_exporter-Liste ≠ blackbox-Liste).
**Rescue-Kanal (universell für LXC auf RBD)**:
```bash
# Auf dem Hyper visor, der den CT hostet:
lsblk # rbdX finden (size matchen)
mount /dev/rbdX /mnt/rescue
# Dateien direkt editieren unter /mnt/rescue/opt/...
umount /mnt/refuge
# im CT: docker compose up -d --force-recreate <svc>
```
**Symptom-Signature Crashloop**: `docker logs prometheus` → "yaml: line N: mapping
values are not allowed in this context" = klassisches NN:-Präfix-Gift.
**Stale-Series-Note**: gelöschte Targets erscheinen sekundenweise weiter in Queries —
erst nach Scrape-Zyklus-Gap re-checken, dann Erfolg erklären.
Ref: docs/solutions/bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md
+22
View File
@@ -0,0 +1,22 @@
# PAT-011 — PVE Webhook Notifications Pipeline Repair
**Trigger**: PVE notifications (esp. vzdump) reach Telegram not / test returns 500.
**Signature diagnosis ladder**:
1. `pvesh create /cluster/notifications/targets/<name>/test` — error taxonomy:
- `Connection refused` → transport/down (check URL target alive)
- `failed to render webhook body` → template broken (syntax/base64)
- `http status: 400` → bridge rejected payload (escaping!)
2. Listener-Sweep über VLAN: `for ip in .xx…; do /dev/tcp/$ip/port probe; done`
3. Byte-Level-Truth via temp mini-sniffer (redirect URL, capture, restore!).
**Hard rules**:
- Mutations an `/etc/pve/notifications.cfg` NUR via `pvesh set` (hand-edits poison
global deserialization!). Vorher Snapshot.
- `body`/header-values/secrets = **base64 blobs**: `printf '%s' tpl | base64 -w0`.
- Handlebars helpers: `{{escape title}}` (Space!), NICHT `{{escape:title}}`.
- Immer `{{escape title}}`/`{{escape message}}` verwenden — rohe Interpolation bricht
bei Apostrophen/Multiline (Backup-Reports!).
- Sniffer-Redirect URL IMMER restaurieren.
Reference: `docs/solutions/bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md`
+49
View File
@@ -0,0 +1,49 @@
---
pattern_id: PAT-010
title: "PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature)"
category: infrastructure
severity: high
status: active
first_observed: 2026-09
last_updated: 2026-09-30
related_systems: [proxmox-cluster, rke2-kubernetes]
related_solution_docs: [docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md]
related_skills: [rke2-cluster-administration, systematic-debugging]
---
# PAT-010: PVE host OOM-kill freezes guest VM as zombie
## Symptom
- K8s node NotReady, kubelet heartbeat stops abruptly
- PVE shows VM `running`, but: no ping, no SSH, qemu-guest-agent dead,
tap-interface TX counters static
- **Signature:** `qm status <vmid> --verbose` → kvm RSS collapses to a
tiny fraction (<5%) of assigned RAM — guest kernel no longer touches
its memory
- Downstream alerts (DS misscheduled, workload CrashLoops) fire, but
NOTHING alarms on the actual OOM
## Root Cause
Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks
the largest anon-RSS process = biggest kvm. After TWO consecutive kills of
the same guest (HA auto-restarts in between), the revived QEMU comes up
but the guest kernel stays inert → zombie VM.
## Mitigation
1. Relieve host pressure FIRST (move movable tenants away) — otherwise
the unfreeze re-boots the guest into the same thrash.
2. Hard cycle the VM: `qm stop` + `qm start` (respect HA guards).
3. Re-query `ha-manager status` afterwards — HA may relocate the VM.
4. Uncordon the K8s node; orphan DS pods self-reconcile.
## Prevention
- Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/
IaC; swap is a buffer, not capacity.
- PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently
invisible to alerting (noticed only via downstream K8s symptoms).
## Evidence
- 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed
rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second
revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration +
hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.
+40
View File
@@ -0,0 +1,40 @@
# Skill Impact Tracker
> Inspired by WikiSkill (arXiv:2608.27454) — tracks skill modifications, their validation
> outcomes, and acceptance decisions. Prevents repeating failed modifications and provides
> an audit trail of skill evolution.
## How to Use
When a skill is patched, created, or deleted, append an entry to the table below.
The entry records WHAT changed, WHY, and WHETHER it helped (if validated).
### Entry Format
```
| Date | Skill | Change Type | Description | Validation | Outcome | Ref |
```
- **Change Type**: `created` | `patched` | `deleted` | `adopted`
- **Validation**: How was success measured? (`manual` | `tested` | `benchmark` | `n/a`)
- **Outcome**: `accepted` | `rejected` | `rolled-back` | `pending`
- **Ref**: Session ID or solution doc path
## Audit Trail
| Date | Skill | Change Type | Description | Validation | Outcome | Ref |
|------|-------|-------------|-------------|------------|---------|-----|
| 2026-08-30 | compound-learning | patched | Added WikiMaintainer phase + patterns/ guidance from WikiSkill paper analysis | tested | accepted | this session |
| 2026-08-30 | (retrospective) | created | 4 solution docs + 2 new patterns (PAT-008, PAT-009) from 30-day retrospective | tested | accepted | this session |
| 2026-08-13 | hindsight | patched | Expanded meta-memory automation table (decay scoring, drift detection, skill health) | tested | accepted | session 2026-08-13 |
| 2026-08-13 | (3 cron jobs) | created | Decay scoring, drift detection, skill health check — no_agent cron jobs | tested | accepted | session 2026-08-13 |
## Notes
- **Rejected proposals persist here** — they are NOT deleted. Future skill updates can
consult this log to avoid repeating failed approaches (the WikiSkill paper showed this
is critical for effective evolution).
- **Neutral changes** (no improvement, no regression) are recorded as `accepted` with
`validation: manual` — they may enable future improvements.
- **Rollbacks** are explicitly tracked — if a skill patch caused regressions and was
reverted, the entry stays with `outcome: rolled-back`.
+58
View File
@@ -0,0 +1,58 @@
---
pattern_id: PAT-002
title: "Traefik `reload` unreliable after conf.d edits — use `restart`"
category: infrastructure
severity: medium
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [proxmox-cluster, rke2-kubernetes]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-26-paperless-ingressroute-hostname-fix.md
related_skills: []
---
# Traefik `reload` unreliable after conf.d edits — use `restart`
## Symptom
After modifying Traefik configuration files (especially dynamic config in `conf.d/`),
issuing a `reload` signal (SIGHUP) does not reliably pick up the changes. The old
configuration remains active, leading to stale ingress routes, incorrect routing,
or 404 errors.
## Root Cause
Traefik's hot-reload mechanism for file-based dynamic configuration can silently fail to
detect changes, especially when:
- Files are edited in-place (atomic rename not used)
- The file watcher misses events on certain filesystems (e.g., overlayfs, NFS)
- rsync is used without `--inplace` (creates temp file + rename, which the watcher may miss)
## Mitigation
Use `restart` instead of `reload` for Traefik container/service after conf.d edits:
```bash
# Instead of: docker kill -s HUP traefik (or systemctl reload traefik)
# Use:
docker compose restart traefik
# or: systemctl restart traefik
```
Additionally, when syncing config files via rsync, use `--inplace` to avoid
temp-file-rename patterns that confuse file watchers:
```bash
rsync --inplace -av ./conf.d/ /etc/traefik/conf.d/
```
## Prevention
- Always use `restart` (not `reload`) after Traefik config changes in the Traefik CT (99999)
- Use `rsync --inplace` when pushing config files to the Traefik host
- Document this in deployment runbooks
## Evidence
- Observed during Paperless v3 IngressRoute hostname fix (Jul 2026)
- Traefik runs in CT99999, conf.d directory
- MEMORY.md entry: "Traefik CT99999: `reload` unreliable after conf.d edits. Use `restart` or `rsync --inplace`."
+71
View File
@@ -0,0 +1,71 @@
---
pattern_id: PAT-004
title: "VFIO GPU passthrough fails — missing softdep + incomplete DRM blacklist"
category: infrastructure
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [proxmox-cluster, rke2-kubernetes]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md
- docs/solutions/bug-fixes/2026-07-21-amd-gpu-passthrough-rombar.md
related_skills: []
---
# VFIO GPU passthrough fails — missing softdep + incomplete DRM blacklist
## Symptom
GPU passthrough to a VM fails intermittently or consistently. Symptoms include:
- `/dev/dri/renderD128` missing in the guest
- `amdgpu` driver not loading in guest
- Kernel BUG in host dmesg
- GPU device visible in `lspci` but not bound to `vfio-pci`
## Root Cause
Two intertwined issues:
1. **Module loading race condition**: Without `softdep amdgpu pre: vfio-pci`, the amdgpu
driver races to bind the GPU before vfio-pci can claim it. Normal bind time ~2.4s;
racing bind time ~302s (or never succeeds).
2. **Incomplete DRM blacklist**: If `drm` and `drm_kms_helper` are not blacklisted, the
kernel's DRM subsystem grabs the GPU before vfio-pci, preventing passthrough.
## Mitigation
Apply to the PVE host's modprobe config:
```bash
# /etc/modprobe.d/blacklist-drm.conf
blacklist drm
blacklist drm_kms_helper
# /etc/modprobe.d/amdgpu-vfio.conf
softdep amdgpu pre: vfio-pci
```
Then rebuild initramfs and reboot:
```bash
update-initramfs -u -k all
reboot
```
Also ensure `rombar=1` in the VM config for AMD GPUs (PCI ID 1002:xxxx):
```
# /etc/pve/qemu-server/<VMID>.conf
hostpci0: 0000:XX:YY.Z,pcie=1,rombar=1,x-vga=1
```
## Prevention
- Always configure `softdep` + DRM blacklist BEFORE attempting GPU passthrough on a new node
- Add to provisioning scripts (Ansible/Tofu) for any node with AMD GPU intended for passthrough
- ROCm 7.x: gfx1150 supported, gfx1036 NOT supported — verify GPU compute ISA before assignment
## Evidence
- Observed on ms-a2-1 (race condition, Jul 2026) and worker-05/VM 102 (kernel BUG, rombar)
- Solution docs: `2026-07-24-vfio-pci-module-loading-race-condition.md`,
`2026-07-21-amd-gpu-passthrough-rombar.md`
- MEMORY.md entry: "GPU:2 active—Immich(CT111/n5pro gfx1150)+Frigate(CT120/px7). ROCm7.14=gfx1150 supported,gfx1036 NOT."
+39 -30
View File
@@ -3,71 +3,80 @@ title: IP-Map (Quick Reference)
category: reference category: reference
tags: [ip, network, reference, quick-lookup] tags: [ip, network, reference, quick-lookup]
created: "2026-07-24" created: "2026-07-24"
modified: "2026-07-24" modified: "2026-09-30"
--- ---
# IP-Map # IP-Map
## Proxmox Hosts (10.0.20.x) — 9 Nodes ## Proxmox Hosts (10.0.20.x) — 8 Nodes
| IP | Hostname | Node ID | Notes | | IP | Hostname | Node ID | Notes |
|----|----------|---------|-------| |----|----------|---------|-------|
| 10.0.20.20 | proxmox2 | 3 | |
| 10.0.20.30 | proxmox3 | 4 | | | 10.0.20.30 | proxmox3 | 4 | |
| 10.0.20.40 | proxmox4 | 2 | | | 10.0.20.40 | proxmox4 | 2 | MON |
| 10.0.20.50 | proxmox5 | 5 | MON, MGR | | 10.0.20.50 | proxmox5 | 5 | MON (leader) |
| 10.0.20.60 | proxmox6 | 6 | | | 10.0.20.60 | proxmox6 | 6 | |
| 10.0.20.70 | proxmox7 | 7 | MON, Traefik CT99999 | | 10.0.20.70 | proxmox7 | 7 | MON, Traefik CT99999 |
| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002 | | 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002, MON, MGR (active) |
| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0 | | 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0, MON |
| 10.0.20.93 | ms-a2-2 | 1 | OSD 14 (SSD 1.8TB), GPU 1002:13c0 | | 10.0.20.93 | ms-a2-2 | 1 | OSD 14 (SSD 1.8TB), GPU 1002:13c0 |
> ⚠️ n5pro (10.0.20.91, nodeid 8) ≠ proxmox7 (10.0.20.70, nodeid 7) — separate Nodes! > ⚠️ n5pro (10.0.20.91, nodeid 8) ≠ proxmox7 (10.0.20.70, nodeid 7) — separate Nodes!
> proxmox1 (10.0.20.10, nodeid 1) wurde dauerhaft entfernt (2026-07-24). > proxmox1 (10.0.20.10) und proxmox2 (10.0.20.20) wurden dauerhaft entfernt.
> Immer `pvecm nodes` für kanonische Liste prüfen. > Immer `pvecm nodes` für kanonische Liste prüfen.
## Kubernetes Nodes (10.0.30.5x-6x) ## Kubernetes Nodes (10.0.30.5x-6x)
| IP | Node | Role | | IP | Node | Role | Status |
|----|------|------| |----|------|------|--------|
| 10.0.30.51 | cp-01 | Control Plane | | 10.0.30.51 | cp-01 | Control Plane | Ready |
| 10.0.30.52 | cp-02 | Control Plane | | 10.0.30.52 | cp-02 | Control Plane | Ready |
| 10.0.30.53 | cp-03 | Control Plane | | 10.0.30.53 | cp-03 | Control Plane | Ready |
| 10.0.30.63 | worker-01 | Worker | | 10.0.30.61 | worker-01 | Worker | Ready |
| 10.0.30.64 | worker-04 | Worker (GPU ✅) | | 10.0.30.64 | worker-04 | Worker (GPU) | **NotReady** |
| 10.0.30.65 | worker-05 | Worker (GPU defekt) | | 10.0.30.65 | worker-05 | Worker (GPU) | Ready |
## Database Layer (10.0.30.7x-8x) ## Database Layer (10.0.30.7x-8x)
| IP | Host | Service | | IP | Host | Service |
|----|------|---------| |----|------|---------|
| 10.0.30.70 | — | MaxScale VIP (keepalived) :3306 | | 10.0.30.70 | — | MaxScale VIP (keepalived) :3306 |
| 10.0.30.71 | VM300 | Galera db1 (ms-a2-1) | | 10.0.30.71 | VM300 | Galera db1 (n5pro) |
| 10.0.30.72 | VM301 | Galera db2 (proxmox3) | | 10.0.30.72 | VM301 | Galera db2 (proxmox6) |
| 10.0.30.73 | VM302 | Galera db3 (proxmox6) | | 10.0.30.73 | VM302 | Galera db3 (ms-a2-2) |
| 10.0.30.81 | VM310 | MaxScale-01 Admin :8989 | | 10.0.30.81 | VM310 | MaxScale-01 Admin :8989 |
| 10.0.30.82 | VM311 | MaxScale-02 Standby | | 10.0.30.82 | VM311 | MaxScale-02 Standby |
## Infrastructure VMs (10.0.30.x) ## Infrastructure VMs (10.0.30.x)
| IP | Host | Service | | IP | Host | Service |
|----|------|---------| |----|------|---------|
| 10.0.30.66 | sarah-hermes (VM107, n5pro) | Sarahs Hermes: WebUI :8787 (LAN-only, pw) + TG-Bot (geplant) |
| 10.0.30.99 | CT111 | Immich (n5pro), Migration zu K8s geplant | | 10.0.30.99 | CT111 | Immich (n5pro), Migration zu K8s geplant |
| 10.0.30.100 | CT134 | InfluxDB (Migration zu K8s) | | 10.0.30.100 | ubuntu | Physischer Ubuntu-Node (Ceph OSDs, ZFS pool01_n2_redundant) |
| 10.0.30.105 | — | ~~Gitea~~ DECOMMISSIONED (CT108, stopped) | | 10.0.30.105 | — | ~~Gitea~~ CT108 GESTOPPT 2026-09-17 (Archive: /home/debian/git-archive/ct108-final/) |
| 10.0.30.124 | VM200 | IaC Runner (DHCP), SSH `debian` | | 10.0.30.124 | VM200 | IaC Runner (DHCP), SSH `debian` |
| 10.0.30.141 | CT141 | Monitoring (Prometheus/Grafana) | | 10.0.30.141 | CT141 | Monitoring (Prometheus/Grafana) |
| 10.0.30.145 | CT145 | Voice Pipeline (PJSIP :5060) | | 10.0.30.145 | CT145 | Voice Pipeline (PJSIP :5060) |
## K8s LoadBalancers (10.0.30.2xx) ## K8s LoadBalancers (10.0.30.2xx) — Cilium LB IPAM, all pinned
| IP | Service | Namespace | | IP | Service | Namespace | Port |
|----|---------|-----------| |----|---------|-----------|------|
| 10.0.30.201 | Hindsight API :9177 | hindsight | | 10.0.30.200 | Gitea SSH | gitea | 22 |
| 10.0.30.204 | InfluxDB :8086 | influxdb | | 10.0.30.201 | Grafana | monitoring | 3000 |
| 10.0.30.205 | Grafana | monitoring | | 10.0.30.202 | Prometheus | monitoring | 9090 |
| 10.0.30.206 | Prometheus :9090 | monitoring | | 10.0.30.203 | Traefik (K8s Ingress) | kube-system | 80/443 |
| 10.0.30.207 | Loki :3100 | logging | | 10.0.30.204 | InfluxDB | influxdb | 8086 |
| 10.0.30.205 | PostgreSQL (CNPG rw LB) | postgres | 5432 |
| 10.0.30.208 | Hindsight API | hindsight | 9177 |
## External CTs (10.0.30.xxx)
| IP | CT/VM | Service | Notes |
|----|-------|---------|-------|
| 10.0.30.99 | CT111 | Seafile v13.0.19 (Docker) | cloud.familie-schoen.com, routed via K8s Traefik IngressRoute |
> All IPs pinned via `io.cilium/lb-ipam-ips` annotation. Pool: 10.0.30.200–250.
## DMZ / Reverse Proxies (10.0.60.x) ## DMZ / Reverse Proxies (10.0.60.x)
| IP | Host | Service | | IP | Host | Service |
|----|------|---------| |----|------|---------|
| 10.0.60.10 | CT9999 | Traefik Reverse Proxy (root/[REDACTED]) | | 10.0.60.10 | CT99999 | Traefik Outer-Proxy (root/[REDACTED]); Terminiert *.familie-schoen.com + leitet schoen.codes als Plain-HTTP an 10.0.30.203 |
## Related ## Related
- [[concepts/network-architecture]] - [[concepts/network-architecture]]
+1 -1
View File
@@ -23,7 +23,7 @@ modified: "2026-07-24"
| 9090 | Prometheus | 10.0.30.206 | LB | | 9090 | Prometheus | 10.0.30.206 | LB |
| 9093 | Alertmanager | K8s internal | | | 9093 | Alertmanager | K8s internal | |
| 9095 | HolmesGPT Adapter | K8s internal | Alertmgr → HolmesGPT | | 9095 | HolmesGPT Adapter | K8s internal | Alertmgr → HolmesGPT |
| 9177 | Hindsight API | 10.0.30.201 | LB, NOT localhost | | 9177 | Hindsight API | 10.0.30.208 | LB, NOT localhost |
| 9221 | PVE Exporter | 10.0.30.141 | | | 9221 | PVE Exporter | 10.0.30.141 | |
| 9283 | Ceph Prometheus | 10.0.20.91 | active mgr | | 9283 | Ceph Prometheus | 10.0.20.91 | active mgr |
| 9345 | RKE2 Server URL | 10.0.30.50 | cp-01 | | 9345 | RKE2 Server URL | 10.0.30.50 | cp-01 |
+47 -52
View File
@@ -3,37 +3,44 @@ title: Ceph Cluster
category: systems category: systems
tags: [ceph, storage, rbd, ec-pool, osd] tags: [ceph, storage, rbd, ec-pool, osd]
created: "2026-07-24" created: "2026-07-24"
modified: "2026-07-25" modified: "2026-09-26"
--- ---
# Ceph Cluster # Ceph Cluster
## Overview ## Overview
- **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9 - **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
- **Health**: HEALTH_OK (recovery complete after OSD 0+2 drain, 0.3% misplaced settling) - **Health**: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed)
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH 2026-07-25), 3 MONs (proxmox5/7/4), MGR on proxmox5 - **Version**: 20.2.4 (tentacle) — all 17 OSDs
- **OSDs**: 13 (8 SSD, 5 HDD), all up/in — OSDs 0+2 destroyed+purged 2026-07-25 - **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1)
- **Capacity**: ~22 TiB total, 6.0 TiB used - **OSDs**: 17 (10 HDD, 7 SSD), all up/in
- **Capacity**: ~33 TiB total, 9.0 TiB used, 24 TiB avail
- **Pools**: 13 pools, 533 PGs (532 active+clean, 1 scrubbing)
## OSD Layout ## OSD Layout
| OSD | Class | Size | Host | Reweight | Notes | | OSD | Class | Size | Host | Reweight | Notes |
|-----|-------|------|------|----------|-------| |-----|-------|------|------|----------|-------|
| 0 | ssd | 188 GB | proxmox2 | — | **DESTROYED 2026-07-25** (92% wear) | | 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | Large HDD | | 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
| 2 | ssd | 233 GB | proxmox2 | — | **DESTROYED 2026-07-25** (slow ops, 81% full) | | 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
| 3 | ssd | 238 GB | proxmox4 | 0.95 | | | 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
| 4 | ssd | 233 GB | proxmox3 | 0.95 | | | 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
| 5 | ssd | 238 GB | proxmox5 | 0.90 | 80% full | | 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
| 6 | hdd | 2.8 TiB | ubuntu | 1.0 | Large HDD | | 7 | hdd | 931 GB | proxmox7 | 0.95 | |
| 7 | hdd | 500 GB | proxmox7 | 1.0 | Was 0.80, reweighted 2026-07-24 | | 8 | hdd | 3.6 TiB | ubuntu | 1.0 | |
| 8 | hdd | 2.8 TiB | ubuntu | 1.0 | BlueFS spillover |
| 9 | ssd | 1.9 TiB | n5pro | 1.0 | | | 9 | ssd | 1.9 TiB | n5pro | 1.0 | |
| 10 | hdd | 300 GB | proxmox6 | 1.0 | Very small HDD | | 10 | hdd | 931 GB | proxmox6 | 0.95 | |
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | Large HDD | | 11 | hdd | 2.8 TiB | n5pro | 1.0 | |
| 12 | ssd | 1.9 TiB | n5pro | 1.0 | | | 12 | ssd | 1.9 TiB | n5pro | 1.0 | |
| 13 | ssd | 1.8 TiB | ms-a2-1 | 1.0 | | | 13 | ssd | 1.8 TiB | ms-a2-1 | 0.95 | |
| 14 | ssd | 1.8 TiB | ms-a2-2 | 1.0 | New 2026-07-24, nvme0n1 | | 14 | ssd | 1.8 TiB | ms-a2-2 | 0.95 | |
| 15 | ssd | 1.8 TiB | ubuntu | 1.0 | New |
| 17 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
| 18 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
> OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD.
> OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded).
## Pools ## Pools
@@ -44,7 +51,7 @@ modified: "2026-07-25"
| 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on | | 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on |
| 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | | | 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | |
| 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on | | 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on |
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true (was 120, equalized to 112) | | 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true |
| 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB | | 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB |
| 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites | | 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites |
| 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | | | 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | |
@@ -58,46 +65,34 @@ modified: "2026-07-25"
## Known Issues ## Known Issues
### Weight Imbalance Causing Placement Failures (2026-07-24) ### HEALTH_WARN: Insecure Key Types (2026-09-26)
HDD hosts have extreme weight disparity: n5pro=10.15TB, ubuntu=5.49TB, proxmox7=0.50TB, proxmox6=0.30TB. Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4.
CRUSH host-level selection (rule 1) often picks only 2 of 4 HDD hosts → up sets with 2 OSDs instead of 3. Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not already set).
Result: PGs stuck in `active+clean+remapped` because up set < min_size.
**Mitigation (2026-07-24)**: ### Small SSDs causing reweightdown
1. Reweighted osd.7 from 0.80 → 1.0 → fixed EC pool 8.3d (NONE → osd.7) osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
2. Equalized pool 6 pg_num 120 → 112 + nopgchange=true Consider removing from CRUSH or replacing with larger drives.
3. Manual pg-upmap for stuck PGs: 5.13 → [1,6,7], 6.6c → [11,8,7], 6.58 → [11,6,7]
4. All `clean+remapped` eliminated. Triggered rebalancing wave (43 PGs backfilling at 26 MiB/s).
**Long-term**: Small HDDs (osd.7 0.5TB, osd.10 0.3TB) cause CRUSH placement failures. Replace with larger disks or create separate CRUSH root for large HDDs only. ### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
thermal throttle T1 3×.**
### Pool 6 pg_num/pgp_num Mismatch (Fixed 2026-07-24) ### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
Pool hdd_disk had pg_num=120, pgp_num=112 (autoscaler reducing to 32). Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
Equalized pg_num to 112. Set nopgchange=true to prevent further autoscaler interference. Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
### BlueFS Spillover on osd.8 ### worker-04 (VM 139) NotReady in K8s
osd.8 spilled 128KiB metadata from db device (2.1GiB of 30GiB) to slow device. Node offline — not a Ceph issue but affects Ceph CSI attachments.
Cosmetic warning, no data risk. Fix: `ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-8`
### Slow Operations on osd.2 and osd.7
osd.2 (81% full, fragmentation 0.80) and osd.7 (small HDD) experience slow BlueStore ops.
osd.2 NVMe has 92% wear — candidate for replacement.
### osd.0 NVMe Wear
92% Wear, Critical Warning → Austausch planen.
### EC Pool k=4+m=1 — No Rebalance Headroom
With 5 OSDs kein Rebalance Headroom. Siehe Solution Doc: `docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`
## RBD Management
- Proxmox RBD Double-Mount Deadlock Pitfall: Niemals `pct mount` und `pct exec` gleichzeitig auf demselben Container
- Siehe Solution Doc: `docs/solutions/bug-fixes/2026-07-23-proxmox-rbd-double-mount-deadlock.md`
## Access ## Access
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.92` - SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50`
- Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs` - Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs`
- Mon nodes: proxmox5, proxmox7, proxmox4 - Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro
- Mgr: proxmox5 (active) - Mgr: n5pro (active)
## Related Skills ## Related Skills
- `ceph-cluster-administration` (devops) - `ceph-cluster-administration` (devops)
+62
View File
@@ -0,0 +1,62 @@
---
title: "Frigate NVR"
category: systems
tags: [frigate, nvr, camera, ai, mqtt, proxmox]
created: "2026-09-27"
modified: "2026-09-27"
---
# Frigate NVR
> Frigate v0.18.0 auf Proxmox LXC CT151. KI-gestützte Objekterkennung für Kameras Einfahrt + Terrasse.
## Infrastruktur
- **Container:** CT151 auf proxmox6
- **IP:** `10.0.30.104:5000`
- **Version:** v0.18.0-
- **Detector:** OpenVINO CPU (3.2 FPS)
- **go2rtc:** v1.9.14
- **GPU:** Intel iGPU `/dev/dri/renderD128` (gid=993)
## Kameras
| Name | IP | Substream | Detect FPS |
|------|----|-----------|------------|
| Einfahrt | `10.0.50.102` | Unterstream | 5.1 fps, 1.0 det |
| Terrasse | `10.0.50.103` | Unterstream | 5.0 fps, 2.2 det |
## MQTT
- **Broker:** `10.0.30.10:1883` (Mosquitto auf HA)
- **Prefix:** `frigate`
- **User:** `frigate`
- **Topic:** `frigate/events`
## Frigate Plus Model
- `plus://709bb8ad097786a7f2a37e27684c79a9`
- Benötigt `PLUS_API_KEY` (Env-Var in `/config/frigate.env`)
## Features
- Face Recognition (groß): Sarah, Dominik, DHL1, Amazon1, Cleo, Lisa, Annemarie, Marcel Schnitzer, Uli, Eva, Postbote, Fr. Schnitzer
- License Plate Recognition (CPU, threshold 0.5)
- Semantic Search (groß)
- Record: alerts+detections, retain 30 days
## HA Automation
- **ID:** `1714065780832` in `/config/automations.yaml`
- **Mode:** `single`, cooldown 600s
- **Dedup:** `input_text.frigate_last_event_id` speichert letzten Event
- **Flow:** MQTT trigger → 10s delay → snapshot → notify → LLM Vision (optional)
- **Labels:** person, car, license_plate, face, dog, cat, amazon, ups, package
- **Sublabel-Extraktion:** `sub[0]` (Jinja, Array-Leck-Fix)
## Bekannte Issues
- Stationäre Autos triggerten wiederholt "new" Events → Flood → gefixt mit cooldown+dedup
- MQTT ACL blockierte HA-Subscription (User `opendtu` hatte keine Subscribe-Rechte) → gefixt
- v0.18 Breaking Changes: `clean_copy` entfernt, `license_plate.mask` Format geändert (list→dict)
## Old Container
- CT120 (frigate): gestoppt, wartet auf Deletion
## Related
- [[systems/homeassistant]] — MQTT, Automations, Notifications
- [[systems/noris-ai]] — LLM Vision für Event-Klassifizierung
- [[entities/infrastructure]] — CT151 auf proxmox6
+5
View File
@@ -35,6 +35,11 @@ modified: "2026-07-24"
- VM300/301/302 (Galera): Fluent Bit aktiv - VM300/301/302 (Galera): Fluent Bit aktiv
- VM310/311 (MaxScale): Fluent Bit aktiv - VM310/311 (MaxScale): Fluent Bit aktiv
## Bekannte Issues
- **2026-09-29:** VM302 Zombie-State nach Migration (QEMU running, Services tot, QGA down). Behoben via Stop+Start; SST-Rejoin ~2-3min. Siehe `docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md`
- **QEMU-Cmdline-Falle:** `-incoming unix:/run/qemu-server/NNN.migrate -S` bleibt nach Migration im Prozess-String — KEIN Beweis für "paused". Immer `qm monitor <vmid> <<< 'info status'` prüfen.
- **HA-Guard:** `ha-manager disable` existiert nicht. Bei Resource-State `request_stop` geht `qm stop` trotzdem durch.
## Related Skills ## Related Skills
- `mariadb-galera-cluster-administration` (devops) - `mariadb-galera-cluster-administration` (devops)
+22 -6
View File
@@ -3,18 +3,18 @@ title: Gitea (Git Server + CI)
category: systems category: systems
tags: [gitea, git, ci, actions] tags: [gitea, git, ci, actions]
created: "2026-07-24" created: "2026-07-24"
modified: "2026-07-26" modified: "2026-09-27"
--- ---
# Gitea (Git Server + CI) # Gitea (Git Server + CI)
## Instanz ## Instanz
- **URL:** git.schoen.codes (K8s Ingress, Traefik → gitea-http ClusterIP) - **URL:** git.schoen.codes (K8s Ingress, Traefik → gitea-http ClusterIP)
- **Internal:** gitea-internal.gitea:3000 (K8s ClusterIP 10.43.5.203) - **Internal:** gitea-http ClusterIP (headless, `None`) — Port 3000/TCP, via Ingress/Traefik erreichbar
- **SSH:** 10.0.30.202:22 (Cilium LoadBalancer, gitea-ssh service) - **SSH:** 10.0.30.200:22 (Cilium LoadBalancer, gitea-ssh service, NodePort 31441) — NICHT .202 (alter Wiki-Fehler)
- **Version:** 1.27.0 (K8s Helm chart v12.7.0) - **Version:** 1.27.0 (K8s Helm chart v12.7.0)
- **Privat:** NIEMALS öffentlich machen (enthält Secrets) - **Privat:** NIEMALS öffentlich machen (enthält Secrets)
- **CT108 (alte Gitea, 10.0.30.105): DECOMMISSIONED** — gestoppt am 2026-07-24 - **CT108 (alte Gitea, 10.0.30.105, v1.25.5): GESTOPPT am 2026-09-17 (genehmigt).** War bis zum Abend NICHT gestoppt trotz älterer Wiki-Claims — Live-Checks (v1.25.5-API-Antwort) schlugen die Papierlage. Vor dem Stop: Full-Census (31 Repos auth, 23 gemeinsam — 22 Tips identisch, dominik/memory K8s-Superset), 29 Bare-Bundle-Archive (143 MB) unter `/home/debian/git-archive/ct108-final/`. Letzter Live-Konsument war die Route `git.familie-schoen.com` im CT99999-Traefik (Upstream .105:3000) — umgebogen auf .203 (K8s), Alias-Host im K8s-Ingress ergänzt (Commit 9e9b6ee). Seither servieren BEIDE Hostnamen v1.27.0.
## Repositories ## Repositories
| Repo | Zweck | Clone | | Repo | Zweck | Clone |
@@ -28,12 +28,15 @@ modified: "2026-07-26"
- Ed25519 SSH-Key als Gitea Deploy Key (read-only) - Ed25519 SSH-Key als Gitea Deploy Key (read-only)
- ArgoCD Secret `argocd-repo-k8s-gitea` mit SSH Private Key - ArgoCD Secret `argocd-repo-k8s-gitea` mit SSH Private Key
- `known_hosts` ConfigMap mit Gitea SSH Hostkey - `known_hosts` ConfigMap mit Gitea SSH Hostkey
- 15 ArgoCD Apps via SSH (`ssh://gitea@git.schoen.codes:22/dominik/iac-homelab.git`) - 17 ArgoCD Apps via SSH (`ssh://git@10.0.30.200:22/dominik/iac-homelab.git`)
- 25 ArgoCD Apps gesamt (Stand Sep 2026)
## Gitea Actions CI ## Gitea Actions CI
- **Primary Runner:** vm200-host (ID 73, act_runner v0.2.13, host backend) — WORKING - **Primary Runner:** vm200-host (ID 73, act_runner v0.2.13, host backend) — WORKING
- **K8s Runner:** gitea-runner deployment in gitea namespace (DinD sidecar + act_runner v0.2.13) — registered but NO TASK ASSIGNMENT (Gitea 1.27.0 bug) - **K8s Runner (gitea NS):** gitea-runner deployment (DinD sidecar + act_runner v0.2.13) — RUNNING, declare erfolgreich (Sep 2026). Vorher 4 Tage Init-crash wegen CSI RBAC Issue (siehe unten).
- **K8s Runner (schoenkitchen NS):** gitea-runner StatefulSet (gleiche Architektur) — RUNNING, declare erfolgreich (Sep 2026).
- **KNOWN BUG (Gitea 1.27.0):** `actions_ready_job` queue doesn't create `action_task` records for K8s-based runners via gRPC Declare. Only runner 73 (external, pre-registered) receives tasks. Fix: upgrade Gitea to 1.27.1+ or 1.28.x. - **KNOWN BUG (Gitea 1.27.0):** `actions_ready_job` queue doesn't create `action_task` records for K8s-based runners via gRPC Declare. Only runner 73 (external, pre-registered) receives tasks. Fix: upgrade Gitea to 1.27.1+ or 1.28.x.
- **WARNUNG (Sep 2026):** `DEFAULT_ACTIONS_URL = https://gitea.com` ist deprecated in Gitea 1.27.0 — fallback zu `github`. Helm Values setzen nur `ENABLED: true`; der Wert kommt vom Chart-Default. Zu fixen via Helm values override: `DEFAULT_ACTIONS_URL: self`.
- **No runner management API** in Gitea 1.27.0 (`/api/v1/admin/runners` → 404). Manage via DB: `gitea.action_runner` table, `is_disabled` column, `deleted` for soft-delete. - **No runner management API** in Gitea 1.27.0 (`/api/v1/admin/runners` → 404). Manage via DB: `gitea.action_runner` table, `is_disabled` column, `deleted` for soft-delete.
- Workflow: `.gitea/workflows/rebuild-infrastructure.yml` - Workflow: `.gitea/workflows/rebuild-infrastructure.yml`
- Achtung: `actions/checkout@v4` versucht github.com zu erreichen — muss gemapped werden - Achtung: `actions/checkout@v4` versucht github.com zu erreichen — muss gemapped werden
@@ -51,6 +54,19 @@ modified: "2026-07-26"
- **Verwendung:** `GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o StrictHostKeyChecking=no" git push ssh://git@10.0.30.202:22/dominik/iac-homelab.git HEAD:main` - **Verwendung:** `GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o StrictHostKeyChecking=no" git push ssh://git@10.0.30.202:22/dominik/iac-homelab.git HEAD:main`
- **Wichtig:** Gitea HTTP (Port 3000) hat keine externe IngressRoute — SSH (Port 22 via LoadBalancer) ist der einzige zuverlässige Weg für Git-Pushes von außerhalb K8s. - **Wichtig:** Gitea HTTP (Port 3000) hat keine externe IngressRoute — SSH (Port 22 via LoadBalancer) ist der einzige zuverlässige Weg für Git-Pushes von außerhalb K8s.
## Known ArgoCD Issues (Sep 2026)
- **`gitea` app:** Sync=Succeeded, Health=Progressing — Ingress hat keine LoadBalancer IP (Traefik reporting gap). Kosmetisch, Funktionalität OK.
- **`gitea-config` app:** OutOfSync — `Deployment gitea-runner` driftet (manuelle Änderung am Image?). Sync würde helfen.
- **`immich` + `immich-config` apps:** OutOfSync — zwei Apps zeigen auf überlappende Pfade (`rendered` vs roh) im iac-homelab Repo. Architektur-Issue, kein Gitea-Problem.
- **`authelia` app:** Health=Progressing — gleiche Ingress-LB-IP Thematik wie gitea.
## Incident 2026-09-27: CSI RBAC → Gitea Pod Init-Crash
- **Symptom:** Gitea Pod 4+ Tage in `Init:0/3`, schoenkitchen runner in `PodInitializing`.
- **Root Cause:** Ceph CSI RBD provisioner `csi-attacher` hatte keine Berechtigung, `csinodes` zu lesen → VolumeAttachment blieb pending → PVC konnte nicht mounten.
- **Fix:** RBAC ClusterRole `ceph-rbd-external-attacher-runner` um `storage.k8s.io/csinodes` Resource ergänzt, provisioner Pods neu gestartet.
- Nach Fix: Alle 24 VolumeAttachments `true`, Gitea Pod Running nach Pod-Delete.
- **Lesson:** Bei CSI-basierten PVs immer zuerst VolumeAttachment Status prüfen, nicht nur PVC Phase.
## Related ## Related
- [[systems/rke2-kubernetes]] - [[systems/rke2-kubernetes]]
- [[concepts/gitops-workflow]] - [[concepts/gitops-workflow]]
+7 -2
View File
@@ -10,8 +10,10 @@ modified: "2026-07-24"
## Deployment ## Deployment
- **Namespace:** hindsight (K8s) - **Namespace:** hindsight (K8s)
- **API:** LoadBalancer `10.0.30.201:9177` (**NICHT localhost**) - **API:** LoadBalancer `10.0.30.208:9177` (**NICHT localhost**, NICHT .201 — dort sitzt seit Rebuild 2026-08-01 Grafana!)
- **Health:** `curl -s http://10.0.30.201:9177/health` - **Health:** `curl -s http://10.0.30.208:9177/health`
- **Port-Mapping:** Svc 9177 → Container 8888; NodePort-Fallback 31577 (jeder Node)
- **LB-IP-Drift-Warnung:** LB-IPs verschoben sich beim Rebuild 2026-08-01 (hindsight .201→.208). Bei Health-Failure IMMER zuerst `kubectl -n hindsight get svc hindsight-api` gegenprüfen, statt Referenz-IP zu vertrauen
- **Backend:** PostgreSQL + pgvector - **Backend:** PostgreSQL + pgvector
- **Config:** `~/.hermes/hindsight/config.json` (api_url gesetzt) - **Config:** `~/.hermes/hindsight/config.json` (api_url gesetzt)
@@ -32,6 +34,9 @@ modified: "2026-07-24"
- **retain_every_n_turns: 2** — reduziert Duplikate - **retain_every_n_turns: 2** — reduziert Duplikate
- Vor Batch-Retain: Health-Check, dann erst retain calls feuern - Vor Batch-Retain: Health-Check, dann erst retain calls feuern
- Meta-Memory Cleanup Cron: 1st of Month 03:00 (3 Dedup Rounds, Archive >6 Monate) - Meta-Memory Cleanup Cron: 1st of Month 03:00 (3 Dedup Rounds, Archive >6 Monate)
- Decay & Importance Scoring: 1st of Month 03:30 (score=access×age_decay, archive score<0.15)
- User Drift Detection: Mondays 04:00 (compare 14d vs 90d baseline)
- Skill Health Check: Mondays 04:30 (scan 187 SKILL.md for issues)
- Nightly Dream Cycle + Weekly Insight Digest - Nightly Dream Cycle + Weekly Insight Digest
## What NOT to store in Hindsight ## What NOT to store in Hindsight
+70
View File
@@ -0,0 +1,70 @@
---
title: "Home Assistant"
category: systems
tags: [homeassistant, smart-home, mqtt, automation]
created: "2026-09-27"
modified: "2026-09-27"
---
# Home Assistant
> Smart Home Zentrale auf Proxmox. Steuerung, Automatisierung und Benachrichtigungen.
## Infrastruktur
- **Host:** `10.0.30.10` (HAOS VM)
- **SSH:** `hassio@10.0.30.10` (PW: 1P Vault "Hermes")
- **URL:** `https://homeassistant.familie-schoen.com`
- **Container:** `homeassistant` (Docker)
## MQTT
- **Broker:** Mosquitto Add-on v7.1.1 auf localhost:1883
- **HA MQTT User:** ehemals `opendtu` (BROKEN — ACL blockiert), gefixt
- **ACL:** `/etc/mosquitto/acl` definiert `user homeassistant` + `user addons`
- **Auth Plugin:** `go-auth.so` (files,http backends)
## Automations
- **File:** `/config/automations.yaml` (35 Automations)
- **Schreibmethode:** SSH → `docker exec homeassistant chmod 666`, danach restore 644
- **Reload:** REST API mit JWT (HS256, signed from `/config/.storage/auth`)
### Frigate Einfahrt Notification (ID 1714065780832)
- MQTT trigger `frigate/events` → filter `type=='new'` + cameras [Einfahrt,Terrasse] + labels [person,car,...]
- Mode: `single`, cooldown 600s
- Dedup: `input_text.frigate_last_event_id`
- Flow: delay 10s → snapshot → notify iPhone → optional LLM Vision
- Siehe [[systems/frigate]]
## Notify Services
| Service | Status |
|---------|--------|
| `notify.mobile_app_iphone_dominik` | ✅ aktiv |
| `notify.mobile_app_sarahs_iphone_app` | verfügbar |
| `notify.mobile_app_ipad_2` | verfügbar |
| `notify.mobile_app_sm_x205` | verfügbar |
## Entitäten
- Kameras: `camera.einfahrt_2`, `camera.terrasse_2` (Suffix `_2` wegen verwaister Integration)
- Motion: `binary_sensor.einfahrt_motion_2`, `binary_sensor.terrasse_motion_2`
- Input Text: `input_text.frigate_last_event_id` (dedup storage, max 255 chars)
## Snapshots
- Gespeichert: `/config/www/snapshots/{camera}_latest.jpg`
- URL: `/local/snapshots/` (HTTP 200, keine Auth)
## JWT Auth
1. `jwt_key` aus `/config/.storage/auth` lesen
2. Client `Hermes_202606` (token id `1444b2c6757b4d66a6f8f8e5899b4e6a`)
3. HS256 signieren, Bearer Header
4. `?return_response=true` für Service-Call Responses
## Integrations
- Frigate (MQTT)
- LLM Vision (noris AI `gemma-4-31b-it`)
- Tibber (Strom)
- Marstek Speicher (VENUS-E, IP 10.0.50.113)
- Various sensors (Xiaomi BLE, etc.)
## Related
- [[systems/frigate]] — NVR Integration
- [[systems/noris-ai]] — LLM Vision Provider
- [[reference/ip-map]] — IP Assignments
+85
View File
@@ -0,0 +1,85 @@
# Laya Decision Model
> 421M param ModernBERT-large Classifier (Apache 2.0) auf CT152, CPU-only.
> Nutze in normalen Sessions fuer schnelle Klassifizierung, Binaerentscheidungen und Pre-Filter.
## Zugang
- **Endpoint:** `POST http://10.0.30.152:8000/predict`
- **Payload:** `{"state": "<text>", "questions": {...}}`
- **Health:** `GET http://10.0.30.152:8000/health`
- **Presets:** `GET http://10.0.30.152:8000/presets/{router|guard|moderation|triage}`
## Entscheidungstypen (Primitives)
| Typ |用途 | Return Fields |
|-----|------|---------------|
| `choice` | Klassifizierung in N Labels | `choice`, `answer_confidence`, `probabilities` |
| `noul` | Ja/Nein mit Wahrscheinlichkeit | `noul` (0..1), `answer_confidence` |
| `score` | Ordinale Bewertung | `score`, `answer_confidence` |
## Verwendung in Sessions
**Praeferieren fuer:**
- Pre-Filter vor teuren LLM-Calls (z.B. "ist diese Email eine Rechnung?" → nur bei "ja" LLM aufrufen)
- Binaerentscheidungen: alert/skip, escalate/ignore, move/keep
- Multi-Kategorie-Klassifizierung mit Confidence
- Gatekeeping: Notification-Suppression, Alert-Filtering
**NICHT geeignet fuer:**
- Textgenerierung / Zusammenfassungen (dafür LLM verwenden)
- Komplexe Reasoning-Tasks
- Embeddings / Semantische Suche (dafür Harrier/Hindsight)
## Example Call
```python
import json, urllib.request
payload = json.dumps({
"state": "Von: amazon.de\nBetreff: Bestellbestätigung #12345",
"questions": {
"kategorie": {
"type": "choice",
"instructions": "Welche Kategorie?",
"criteria": {
"rechnung": "Rechnung, Invoice",
"bestellung": "Bestellbestätigung, Order",
"werbung": "Newsletter, Marketing"
}
},
"ignorieren": {
"type": "noul",
"instructions": "Soll ignoriert werden?"
}
}
}).encode()
req = urllib.request.Request(
"http://10.0.30.152:8000/predict",
data=payload,
headers={"Content-Type": "application/json"}
)
result = json.loads(urllib.request.urlopen(req, timeout=30).read())
# result["answers"]["kategorie"]["choice"] → "bestellung"
# result["answers"]["kategorie"]["answer_confidence"] → 0.99
```
## Performance
- Latenz: ~1.5-1.7s warm cache (CPU)
- Load time: ~18s (cold start)
- systemd service: `laya.service` (enabled, onboot)
## Einsatzgebiete (aktiv)
- **Rechnungen-Organizer** (Cron `f773f8c23230`): 12-Kategorie Email-Klassifizierung, 3-Schichten-Safety
- Weitere Kandidaten: Beleg-Sammler, SRE Network Recon, Backup Digest, Frigate Event Gate
## Constraints
- 421M Modell → niedrige Confidence bei ambiguous Inputs (Feature, nicht Bug!)
- Batch funktioniert nicht — ein Request pro Input-Instanz
- `noul` ist nicht zuverlaessig fuer kritische Entscheidungen allein — immer mit `choice` kombinieren
- Confidence-Threshold empfohlen (≥0.6 fuer Moves, ≥0.5 fuer Ignores)
## Related
- [[concepts/email-organization]] — Rechnungs-Organizer Architecture
- [[entities/infrastructure]] — CT152 auf proxmox6
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
+40
View File
@@ -0,0 +1,40 @@
---
title: "noris AI Platform"
category: systems
tags: [ai, llm, noris, gpu, embeddings]
created: "2026-09-27"
modified: "2026-09-29"
---
# noris AI Platform (ai.noris.de)
> Interne AI-Plattform der noris Network AG. Bereitstellung von LLMs, Embeddings und Image Generation.
## Endpoints
- **Chat:** `https://ai.noris.de/v1/chat/completions`
- **Embeddings:** `https://ai.noris.de/v1/embeddings`
- **Images:** ⚠️ `/v1/images/generations` wird vom Bifrost Gateway **NICHT** unterstützt. Image-Gen-Modelle (qwen-image-2-1) werden über `/v1/chat/completions` angesprochen — das Bild kommt als base64-PNG im `content`-Array zurück (Typ `image_url`, `data:image/png;base64,...`). Siehe `references/vllm-image-generation.md` im Skill `serving-llms-vllm`.
## Modelle
| Typ | Modell-ID | Hinweise |
|-----|-----------|----------|
| Flagship LLM | `glm-5-2` | Primary, OpenRouter-kompatibel |
| General | `gemma-4-31b-it` | Vision-fähig, genutzt von LLM Vision |
| Large MoE | `gpt-oss-120b` | |
| Mid-range | `qwen3.6-27b` | |
| Mid-range | `qwen3.8-27b` | |
| Fast | `ds-v4-flash` | Low-latency, Paperless OCR |
| Embedding | `harrier` | Vektorembeddings |
| Image Gen | `qwen-image-2-1` | Via `/v1/chat/completions` (NOT images/generations). Base64-PNG im content-Array. ~30s/ Bild. |
## Verbraucher
- **Hermes Agent** — Primärmodell `glm-5-2` via OpenRouter
- **HA LLM Vision** — `gemma-4-31b-it` für Bildanalyse (Frigate Events)
- **Paperless** — `ds-v4-flash` für OCR/Kategorisierung
- **Personal Coach Bot** — `glm-5-2` via noris direkt
- **Dynamic Coach** — `glm-5-2` via noris direkt, `qwen-image-2-1` für Visualisierungen
## Related
- [[systems/frigate]] — nutzt noris AI für Event-Klassifizierung
- [[systems/homeassistant]] — LLM Vision Integration
- [[systems/paperless]] — OCR via ds-v4-flash
+34
View File
@@ -0,0 +1,34 @@
---
title: "Paperless-ngx"
category: systems
tags: [paperless, documents, oidc, ocr]
created: "2026-09-27"
modified: "2026-09-27"
---
# Paperless-ngx
> Dokumentenmanagement mit OCR, OIDC-Login und AI-Kategorisierung.
## Zugriff
- **URL:** `https://dokumente.familie-schoen.com`
- **mTLS:** `https://dokumente-mtls.familie-schoen.com` (auto-login als `dominik`)
- **PKCS12:** `dominik-dokumente-mtls.p12` (PW: siehe 1P Vault "Hermes")
## Auth
- OIDC via Authelia (`dominik@schoen.eu`)
- Break-Glass lokaler User: `dominik`
- mTLS Client Cert: CN=dominik, gültig bis Juli 2028
## Konfiguration
- **AI Backend:** `ds-v4-flash@ai.noris.de` (noris AI)
- **Memory:** 4Gi (PAT-002)
- **Mail Import:** `dokumente@familie-schoen.com` (iCloud mailbox, max 30 Tage)
- **Owner:** dominik
## Known Issue
- PAT-002: Memory-Limit 4Gi erforderlich, sonst OOM bei großen OCR-Batches
## Related
- [[concepts/credential-policy]] — 1Password, mTLS Zertifikate
- [[systems/noris-ai]] — ds-v4-flash für OCR
+87 -8
View File
@@ -3,14 +3,14 @@ title: Proxmox VE Cluster
category: systems category: systems
tags: [proxmox, virtualization, lxc, qemu, pve] tags: [proxmox, virtualization, lxc, qemu, pve]
created: "2026-04-28" created: "2026-04-28"
modified: "2026-07-24" modified: "2026-09-30"
--- ---
# Proxmox VE Cluster # Proxmox VE Cluster
## Cluster-Konfiguration ## Cluster-Konfiguration
- **Version:** PVE 9.2.3 - **Version:** PVE 9.2.20, Kernel 7.0.14-19-pve (upgraded 2026-09-25)
- **Nodes:** 9 (Quorum OK) - **Nodes:** 8 (Quorum OK, proxmox2 dauerhaft entfernt)
- **Hypervisoren:** 10.0.20.x - **Hypervisoren:** 10.0.20.x
- **Guests:** ~30 LXC + ~10 QEMU VMs - **Guests:** ~30 LXC + ~10 QEMU VMs
@@ -50,13 +50,92 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
- Benötigte modprobe.d Config: - Benötigte modprobe.d Config:
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper` - `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci` - `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
- Worker-05 (VM 102) läuft auf ms-a2-2 mit funktionierendem GPU-Passthrough - Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
## Bekannte Probleme ## Bekannte Probleme
- osd.0 NVMe 92% Wear — Austausch planen - CT110 kaputte libc — Reparatur ausstehend (still stopped)
- osd.2/5 nearfull (93-94%) — entlasten - osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
- CT110 kaputte libc — Reparatur ausstehend - osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
- ms-a2-1 GPU-Passthrough: Config gefixt, Reboot zur Verifikation ausstehend - **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
## HA Rules (PVE 9.2 Rules System)
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
**Resource-Affinity (negative = anti-collocation):**
*RKE2 Control Plane (etcd-Quorum braucht 2/3):*
- `rke2-cp-anti-112-122`: vm:112 ↔ vm:122 — nie auf gleichem Host
- `rke2-cp-anti-112-126`: vm:112 ↔ vm:126 — nie auf gleichem Host
- `rke2-cp-anti-122-126`: vm:122 ↔ vm:126 — nie auf gleichem Host
*RKE2 Worker:*
- `rke2-worker-anti-102-128`: vm:102 ↔ vm:128 — nie auf gleichem Host
- `rke2-worker-anti-102-139`: vm:102 ↔ vm:139 — nie auf gleichem Host
- `rke2-worker-anti-128-139`: vm:128 ↔ vm:139 — nie auf gleichem Host
*Galera:*
- `galera-anti-300-301`: vm:300 ↔ vm:301 — nie auf gleichem Host
- `galera-anti-300-302`: vm:300 ↔ vm:302 — nie auf gleichem Host
- `galera-anti-301-302`: vm:301 ↔ vm:302 — nie auf gleichem Host
*MaxScale:*
- `maxscale-anti-310-311`: vm:310 ↔ vm:311 — nie auf gleichem Host
**Node-Affinity (non-strict, failover allowed):**
*RKE2 CP:*
- `na-vm112`: vm:112 → proxmox3, proxmox5, proxmox7
- `na-vm122`: vm:122 → ms-a2-1, ms-a2-2
- `na-vm126`: vm:126 → proxmox4, proxmox5, proxmox6
*RKE2 Worker:*
- `na-vm102`: vm:102 → proxmox5, proxmox6, proxmox7 (NICHT ms-a2-2!)
- `na-vm128`: vm:128 → ms-a2-2, proxmox5, proxmox7
- `na-vm139`: vm:139 → n5pro, proxmox3, proxmox4
*Hermes:*
- `na-vm230`: vm:230 → n5pro, proxmox5, proxmox6
*Galera/MaxScale:*
- `na-vm300`: vm:300 → n5pro, proxmox3, proxmox4
- `na-vm301`: vm:301 → proxmox6, proxmox5, proxmox7
- `na-vm302`: vm:302 → ms-a2-2, ms-a2-1
- `na-vm310`: vm:310 → proxmox7, proxmox5, proxmox4
- `na-vm311`: vm:311 → ms-a2-1, ms-a2-2
**Aktuelle Verteilung (alle Anti-Collocation erfüllt):**
| Role | VM | Node |
|------|----|------|
| RKE2 CP-01 | 112 | proxmox3 |
| RKE2 CP-02 | 122 | ms-a2-1 |
| RKE2 CP-03 | 126 | proxmox4 |
| RKE2 Worker-01 | 128 | proxmox5 |
| RKE2 Worker-04 | 139 | n5pro |
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
| Hermes-Agent-01 | 230 | n5pro |
| Galera db1 | 300 | n5pro |
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
| Galera db3 | 302 | ms-a2-2 |
| MaxScale-01 | 310 | proxmox7 |
| MaxScale-02 | 311 | ms-a2-1 |
> ⚠️ PVE 9.2 Constraints:
> - Resources in resource-affinity rules dürfen keine multi-priority node-affinity haben (gleiche Priorität für alle Nodes erforderlich).
> - `ha-manager add` MUSS vor `ha-manager rules add` kommen — sonst "cannot use unmanaged resource".
> - Bei gleichzeitigem HA-Add + Anti-Collocation-Violation kann HA-Manager deadlocks (beide VMs auf `migrate` fest). Lösung: eine VM temporär aus HA entfernen, manuell migrieren, dann re-add.
> - Online-Migration von VMs mit hohen Memory-Writes (>12GB dirty pages) kann `broken pipe` fehlschlagen. Offline-Migration (stop→migrate→start) als Fallback.
## Related ## Related
- [[systems/ceph-cluster]] - [[systems/ceph-cluster]]
+24 -6
View File
@@ -3,7 +3,7 @@ title: RKE2 Kubernetes Cluster
category: systems category: systems
tags: [kubernetes, rke2, cilium, argocd, cnpg, gitops] tags: [kubernetes, rke2, cilium, argocd, cnpg, gitops]
created: "2026-07-24" created: "2026-07-24"
modified: "2026-07-24" modified: "2026-09-17"
--- ---
# RKE2 Kubernetes Cluster # RKE2 Kubernetes Cluster
@@ -22,16 +22,21 @@ modified: "2026-07-24"
| cp-03 | 10.0.30.53 | Control Plane | | cp-03 | 10.0.30.53 | Control Plane |
| worker-01 | 10.0.30.63 | Worker | | worker-01 | 10.0.30.63 | Worker |
| worker-04 | 10.0.30.64 | Worker | | worker-04 | 10.0.30.64 | Worker |
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) | | worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2; 30.09. OOM-Zombie-Freeze behoben, siehe PAT-010) |
## Storage ## Storage
- **Ceph CSI**: ceph-flash (fast), ceph-hdd (bulk) - **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432` - **⚠️ StorageClass IaC (2026-09-29):** `ceph-flash` + `ceph-hdd-replica` waren MANUELL erstellt (nicht in Git). Jetzt committed unter `clusters/main/storage/`. Alle RBD SCs MÜSSEN `controller-expand-secret-name/namespace` haben — fehlt dies, schlagen Volume Expansions fehl ("provided secret is empty") und triggern Retry-Loops die API Server überlasten.
- **⚠️ PVs snapshotten SC-Parameter zur Provisionierungszeit** — SC fixen reicht NICHT; bestehende PVs brauchen individuellen Patch mit `controllerExpandSecretRef` falls sie ohne erstellt wurden.
- **⚠️ StorageClass `parameters` sind IMMUTABLE** — können nicht gepatched werden, müssen gelöscht+neu erstellt werden.
- **⚠️ Snapshot-CRD-Flavor (seit Rebuild 2026-08-01):** `volumesnapshotclasses`-CRD (RKE2-Addon `rke2-snapshot-controller-crd`) hat **FLAT-Schema** — `driver`/`deletionPolicy`/`parameters` auf TOP-LEVEL, `spec:` existiert nicht im Schema (Upstream-CRD wäre nested!). Niemals struktur-intuitiv "reparieren" — CRD-Schema lesen. Live-VSCs: `ceph-rbd-snapclass` (default) + `cephfs-snapclass`
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`; Specs **100Gi data / 20Gi WAL** (grow-only — Shrink wird von CNPG-Admission verboten, Git immer nach oben alignieren)
- **⚠️ etcd auf Ceph RBD (~20ms WAL fsync auf allen CP-Nodes)** — architektonisches Risiko; unter API Server Write Pressure → gRPC DeadlineExceeded Cascade. Lokales NVMe für etcd data dirs empfohlen.
## GitOps ## GitOps
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea - **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
- 15 Applications via SSH (`ssh://git@10.0.30.202:22/dominik/iac-homelab.git`) - **Feed-Quelle (Stand 2026-09-17):** `http://10.0.30.105:3000/...` = Legacy-CT108-Gitea (noch aktiv!). Geplanter Flip auf `ssh://git@10.0.30.200:22/dominik/iac-homelab.git` (end-to-end verifiziert), danach CT108-Stop (Freigabe Dominik)
- **Gitea SSH user is `git`, not `gitea`** — `git.schoen.codes` resolves to Traefik (10.0.30.203, HTTP only), SSH is on 10.0.30.202:22 - **Gitea SSH user is `git`, not `gitea`** — SSH-LoadBalancer = **10.0.30.200:22** (svc gitea-ssh, targetPort 2222, NodePort 31441). Legacy-.202 ist TOD (Timeout). Repo-Historien am 2026-09-17 konsolidiert (identische Tips auf beiden Remotes, Merge `d76d3e8` + `8c741c5`)
- Workflow: IaC repo auschecken → ändern → commit+push → ArgoCD sync → verify → lokale Kopie löschen - Workflow: IaC repo auschecken → ändern → commit+push → ArgoCD sync → verify → lokale Kopie löschen
- Siehe [[concepts/gitops-workflow]] - Siehe [[concepts/gitops-workflow]]
@@ -61,6 +66,19 @@ modified: "2026-07-24"
- IaC: unified `amd-gpu` PCI mapping (n5pro + ms-a2-1 + ms-a2-2) + dynamic hostpci (one worker block, gpu flag) - IaC: unified `amd-gpu` PCI mapping (n5pro + ms-a2-1 + ms-a2-2) + dynamic hostpci (one worker block, gpu flag)
- ms-a2-1 has kernel BUG with 1002:13c0 (renderD128 missing) — worker-05 moved to ms-a2-2 (identical hardware) - ms-a2-1 has kernel BUG with 1002:13c0 (renderD128 missing) — worker-05 moved to ms-a2-2 (identical hardware)
## Capacity Management
- **Descheduler** v0.30.0 deployed als ArgoCD-App (`clusters/main/descheduler/manifests.yaml`)
- Plugin: `LowNodeUtilization` (thresholds 30/50/20, targetThresholds 70/70/60)
- Interval: 2m (`--descheduling-interval=2m`)
- `nodeFit: true` auf Profile-Ebene (nicht Plugin-Ebene — v0.30 Schema!)
- RBAC: benötigt `list/watch` auf `namespaces` (fehlt in Default-Helm-Chart)
- Raw Manifests statt Helm Chart (Chart 0.30 hat params/args-Bug, Chart 0.35 hat falsche API-Version)
- Siehe Solution Doc `2026-09-30-k8s-cp01-relief-descheduler.md`
- **CP-01 Relief (2026-09-30):** hindsight-api + hindsight-postgres + paperless von cp-01 → worker-05 migriert
- cp-01 RAM: 93% → 43%; worker-05 bei ~28%
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren
- Live-only nodeSelector (hindsight-api, Sep-17-Hotfix) war nie in Git → Live-Patch nötig
## Known Pitfalls ## Known Pitfalls
- `enableServiceLinks: false` bei Apps deren Service-Name mit Env-Vars kollidiert (z.B. Paperless `PAPERLESS_PORT`) - `enableServiceLinks: false` bei Apps deren Service-Name mit Env-Vars kollidiert (z.B. Paperless `PAPERLESS_PORT`)
- ArgoCD `--force` kann nicht mit ServerSideApply kombiniert werden - ArgoCD `--force` kann nicht mit ServerSideApply kombiniert werden
+74
View File
@@ -0,0 +1,74 @@
---
title: Sarah-Hermes (VM107)
category: systems
tags: [hermes, webui, vm, sarah, telegram, backup]
created: "2026-09-30"
modified: "2026-09-30"
related: [systems/rke2-kubernetes, reference/ip-map, systems/noris-ai]
---
# Sarah-Hermes (VM107)
Zweite, vollständig isolierte Hermes-Instanz für Sarah (Allround-Assistentin:
Erinnerungen, Planung, Smalltalk). Getrennte Memories/Sessions/Skills/Keys von
den 5 Owner-Profilen.
## Eckdaten
| Attribut | Wert |
|----------|------|
| VM-ID | 107 (auto-allokiert) |
| Node | n5pro (Template 9000 debian-12-cloudinit) |
| IP | 10.0.30.66/24 (static via DHCP reservation) |
| Specs | 2 vCPU / 4 GB RAM / 32 GB Disk (vm_disks/RBD) |
| Access | SSH `debian@10.0.30.66` mit `~/.ssh/id_ed25519_cloudinit` |
| Tofu | `iac-homelab/epic-8-sarah-hermes/tofu/` |
| Ansible | `iac-homelab/epic-8-sarah-hermes/ansible/` |
| Plan | `iac-homelab/docs/plans/2026-09-30-sarah-hermes-vm.md` |
## Stack
- **hermes-webui** (ghcr.io/nesquena/hermes-webui:latest), Single-Container,
Port 8787 (0.0.0.0 gebunden, UFW erlaubt nur LAN), Password-Auth.
Compose: `/home/debian/hermes-webui/docker-compose.yml` auf der VM.
- **Agent-Runtime:** nousresearch/hermes-agent geklont nach
`~/.hermes/hermes-agent` auf der VM; installiert in `/app/venv` im Container
(editable). **Wichtig:** Core-Deps im pyproject sind hinter
`python_version >= '3.14'`-Markern gepinnt → auf Python 3.12 installiert
`-e .` NULL Deps. Manual-Dep-Bootstrapping nötig (siehe Pitfalls).
- **LLM:** noris-Provider (`https://ai.noris.de/v1`, Default-Modell
`vllm/release/glm-5-2`), Key via `HERMES_CUSTOM_NORIS_API_KEY` aus
`.env` (Compose mapped explizit in den Container).
- **Telegram:** geplant (blockiert auf BotFather-Token von Sarah/Dominik).
Bei Aktivierung: `gateway.telegram_enabled: true` + Token in `.env`;
Webhook/API-Server-Ports bleiben disabled (Konfliktvermeidung).
## Secrets (1Password, Vault: Hermes)
- `sarah-hermes-webui` → HERMES_WEBUI_PASSWORD
- `sarah-hermes-noris-key` → HERMES_CUSTOM_NORIS_API_KEY
## Backup
- Daily vzdump-Job `backup-2e8a34e3-66cb` (23:00, `all=1`, exclude 301,302,310,311,137,147,501)
→ **deckt VM107 ab** (Storage `noris_v4` = PBS Datastore `noris` @ 10.0.30.119).
- Manueller Verify-Lauf am 30.09.: TASK OK in 48s (inkrementell, 88% reuse).
- Offsite: folgt dem regulären `push-offsite` Sync-Job (Pull↔Push-Korrektur
vom 19.09.).
## Known Issues / Pitfalls
1. **Python-Version-Mismatch:** hermes-agent pyproject pins Core-Deps an
`python_version >= '3.14'`; WebUI-Container läuft auf 3.12 → `-e .`
installiert keine Deps. Fix: manuell `pip install` der gepinschten Pakete
+ iterativer Missing-Import-Loop bis `import run_agent` klappt.
2. **hermes update Ownership-Konflikt:** `hermes update`-Completion beschwert
sich über uid 0 vs. uid 1000 auf `/app/venv/bin/hermes-acp`. Kosmetisch,
Betriebsbetrieb unbeeinträchtigt. Fix-Idee: `chown` im Entry-Point.
3. **telegram-bridge Notify-Target defekt** (seit 19.09., Connection refused):
betrifft vzdump-Notifications clusterweit, nicht nur VM107. Separater Fix.
## Verification History
- 2026-09-30: Deploy + E2E-Test (Login 200, Agent-Antwort "HALLO" via glm-5-2).
- 2026-09-30: Manueller vzdump → TASK OK (48s, inkrementell).