Compare commits
14
Commits
7c36a7600a
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b527e4ff3c | ||
|
|
3a57db957e | ||
|
|
cc2f237ca3 | ||
|
|
531d6bfbdd | ||
|
|
a68665eca0 | ||
|
|
d487b596d7 | ||
|
|
f09c6ef414 | ||
|
|
dfd3add11b | ||
|
|
d1464eb9e3 | ||
|
|
02f6a8e2b0 | ||
|
|
f3039765ed | ||
|
|
1d036dbcdc | ||
|
|
fb47c2a182 | ||
|
|
b2e9512a0e |
@@ -14,7 +14,7 @@ Infra-Changes **bevorzugt** über GitOps. Geht nicht immer. Ausnahmen müssen vo
|
||||
## Workflow
|
||||
1. IaC Repo lokal auschecken (`/root/iac-homelab` auf VM200)
|
||||
2. Ändern (Tofu/Ansible/K8s Manifeste)
|
||||
3. Commit + Push (zu BEIDEN Gitea Remotes: origin + k8s)
|
||||
3. Commit + Push zum EINEN kanonischen Remote: `ssh://git@10.0.30.202:22/dominik/iac-homelab.git` (Dual-Push-Habit führte 09/26 zur Ghost-Regression über CT108 — NIEMALS wieder zweites Remote pflegen)
|
||||
4. ArgoCD sync (oder auto-sync)
|
||||
5. Verify (kubectl get, curl, etc.)
|
||||
6. Lokale Kopie löschen
|
||||
|
||||
@@ -0,0 +1,135 @@
|
||||
---
|
||||
title: "Memory Layer Architecture"
|
||||
category: concepts
|
||||
tags: [memory, architecture, hindsight, lcm, wiki, separation]
|
||||
created: "2026-09-27"
|
||||
modified: "2026-09-27"
|
||||
---
|
||||
|
||||
# Memory Layer Architecture — Practical Rules
|
||||
|
||||
## Übersicht
|
||||
|
||||
7 Schichten mit klaren, nicht-überlappenden Rollen.
|
||||
|
||||
| # | Layer | System | Pfad/Ort | Rolle | Wann nutzen |
|
||||
|---|-------|--------|----------|-------|-------------|
|
||||
| L0 | Hot | MEMORY.md + USER.md | `~/.hermes/memories/` | System-Prompt Injection, komprimierte Facts | Jede Session (automatisch) |
|
||||
| L1 | Curated | LLM Wiki | `~/.hermes/memory/` | Browsbares Wissen, git-synced | Wenn Details gebraucht werden |
|
||||
| L2 | Semantic | Hindsight | K8s PostgreSQL | Vektor-Suche, semantisches Retrieval | Cross-Session Lookup |
|
||||
| L3 | Procedural | Skills | `~/.hermes/skills/` | Wie-geht-es Prozeduren mit Pitfalls | Bei wiederkehrenden Tasks |
|
||||
| L4 | Episodic | LCM + session_search | SQLite lcm.db / state.db | Rohe Gesprächsverläufe, Compaction | Innerhalb aktiver Session |
|
||||
| L5 | Solutions | Solution Docs | `~/docs/solutions/` | Problem-Lösungs-Paare | Nach substanziellen Tasks |
|
||||
| L6 | Project | AGENTS.md | pro Repo | Repo-spezifische Konventionen | Beim Arbeiten in einem Repo |
|
||||
|
||||
## Was wohin gehört — Entscheidungsbaum
|
||||
|
||||
```
|
||||
Ist es ein Fact über Dominik persönlich?
|
||||
→ USER.md (L0)
|
||||
Ist es sein Kommunikationsstil, Safety-Rule, oder Workflow-Präferenz?
|
||||
→ USER.md
|
||||
Ist es ein persönlicher Umstand (Größe, Gewicht, Arbeitgeber)?
|
||||
→ Hindsight (L2) — NICHT USER.md
|
||||
|
||||
Ist es ein Infra-Fakt (IP, Version, Konfiguration)?
|
||||
→ LLM Wiki (L1) — systems/ oder reference/
|
||||
Braucht er jede Session sofort?
|
||||
→ MEMORY.md (L0) als 1-Zeiliger Pointer "→Wiki systems/xyz"
|
||||
|
||||
Ist es ein "Wie mache ich X"-Prozess?
|
||||
→ Skills (L3)
|
||||
|
||||
Ist es ein "Wir hatten Problem Y, Lösung war Z"-Eintrag?
|
||||
→ Solution Doc (L5) + Hindsight-Index (L2)
|
||||
|
||||
Ist es ein vergangenes Gesprächsereignis?
|
||||
→ LCM (L4) kümmert sich automatisch darum
|
||||
```
|
||||
|
||||
## L0 — MEMORY.md vs USER.md
|
||||
|
||||
### USER.md (1.375 chars max)
|
||||
**Nur:** Dominiks Persönliche Präferenzen, Safety-Rules, Kommunikationsstil
|
||||
- "NIEMALS ohne Approval löschen"
|
||||
- "Action-first, kein Dom"
|
||||
- "Crash+Alert statt Silent Skip"
|
||||
- "Laya in Sessions nutzen"
|
||||
|
||||
**NICHT:** Infra-Fakten, IP-Adressen, Versionsnummern, Team-Struktur
|
||||
|
||||
### MEMORY.md (2.200 chars max)
|
||||
**Nur:** Komprimierte Pointer auf Wiki-Seiten + nicht-wikiffähige Quick-Facts
|
||||
- `Frigate v0.18 CT151. →Wiki systems/frigate`
|
||||
- `Galera:VM300/301/302,VIP .70. →Wiki systems/galera-maxscale`
|
||||
- Baufinanz-Zahlen (persönlich, nicht im Wiki)
|
||||
- LinkedIn Persona (Workflow, nicht im Wiki)
|
||||
|
||||
**NICHT:** Vollständige Specs, Konfigurationsdetails, lange Beschreibungen
|
||||
|
||||
## L1 — LLM Wiki
|
||||
|
||||
### Wann erstellen/aktualisieren?
|
||||
- Bei jeder infrastrukturellen Änderung (Compound-Learning Cycle Phase 4.7)
|
||||
- Neue Seite wenn: neues System, neuer Service, neue architektonische Entscheidung
|
||||
- Patch wenn: Versionswechsel, IP-Änderung, Known Issue hinzugekommen
|
||||
|
||||
### Wann NICHT?
|
||||
- Procedures → Skills (L3)
|
||||
- Problem-Solution Pairs → Solution Docs (L5)
|
||||
- Credentials → 1Password
|
||||
|
||||
## L2 — Hindsight
|
||||
|
||||
### Was gehört rein?
|
||||
- **Semantische Pointer** auf Solution Docs und wichtige Meilensteine
|
||||
- **Cross-Session Kontext** der nicht im Wiki steht (z.B. "wir haben am Datum X entschieden...")
|
||||
- **Episodisches Wissen** über Projekte, Entscheidungen, Learnings
|
||||
|
||||
### Was NICHT rein gehört?
|
||||
- ❌ Infra-Topologie (→ L1 Wiki)
|
||||
- ❌ Aktuelle IPs/Versionsnummern (→ L1 Wiki, werden schnell stale)
|
||||
- ❌ Procedures (→ L3 Skills)
|
||||
- ❌ Credentials (→ 1Password)
|
||||
- ❌ Session-Fortschritt (→ L4 LCM)
|
||||
|
||||
### Bekanntes Problem: Hindsight Accumulation
|
||||
Hindsight hat über Monate Infra-Fakten angesammelt die inzwischen teilweise
|
||||
veraltet sind (9 vs 8 Nodes, alte OSD-Counts). Hindsight bietet keine
|
||||
Delete-Funktion via Tools. Strategie:
|
||||
1. Zukünftig nur noch semantische Pointer + Entscheidungen speichern
|
||||
2. Topologie-Fakten nicht mehr via `hindsight_retain` ablegen
|
||||
3. Veraltete Einträge ignorieren — Hindsight Ranking bevorzugt neuere Memories
|
||||
|
||||
## L4 — LCM (Local Conversation Memory)
|
||||
|
||||
### Rolle
|
||||
- **Innerhalb aktiver Session:** Compaction, Summary-DAG, Fresh Tail
|
||||
- **Cross-Session (lcm.db):** Rohe Nachrichten + Summary Nodes aller Sessions
|
||||
- **Kein manuelles Management nötig** — läuft automatisch
|
||||
|
||||
### Wann manuell eingreifen?
|
||||
- `lcm_status` bei Verdacht auf Compression-Problemen
|
||||
- `lcm_grep` um exakte Phrasen in vergangenen Sessions zu finden
|
||||
- `lcm_recall` für bedeutungsbasierte Cross-Session-Suche
|
||||
- Niemals manuell Daten in LCM schreiben — es ist Auto-Managed
|
||||
|
||||
## L5 — Solution Docs
|
||||
|
||||
### Wann erstellen?
|
||||
- Nach substanziellem Bugfix (≥5 Tool-Calls)
|
||||
- Nach Architektur-Änderung
|
||||
- Nach komplexer Migration
|
||||
- Nach schwierigem Troubleshooting
|
||||
|
||||
### Format
|
||||
`~/docs/solutions/{type}/{YYYY-MM-DD}-{slug}.md`
|
||||
Types: `architecture/`, `bugfix/`, `migration/`, `workflow/`
|
||||
|
||||
Immer: Hindsight-Index (`hindsight_retain`) nach Erstellung
|
||||
|
||||
## Related
|
||||
- [[systems/hindsight]] — Hindsight Setup, API
|
||||
- [[systems/laya]] — Laya Classifier
|
||||
- Skill: `memory-sync` — Wiki Maintenance Protocol
|
||||
- Skill: `software-development/compound-learning` — Phase 4.7 Wiki Update
|
||||
@@ -1,23 +1,23 @@
|
||||
# User Drift Report — 2026-08-31
|
||||
# User Drift Report — 2026-09-28
|
||||
|
||||
**Recent window:** Last 14 days (382 msgs)
|
||||
**Baseline:** Previous 90 days (5490 msgs)
|
||||
**Recent window:** Last 14 days (1884 msgs)
|
||||
**Baseline:** Previous 90 days (6185 msgs)
|
||||
|
||||
## Detected Drifts
|
||||
|
||||
- **message_length**: Messages 41% longer (baseline: 1083 chars → recent: 1531)
|
||||
- **new_focus**: New dominant topics: ai_ml, galera, security
|
||||
- **declining_focus**: Topics fading from focus: gitops, ceph, k8s
|
||||
- **message_length**: Messages 98% longer (baseline: 1192 chars → recent: 2361)
|
||||
- **new_focus**: New dominant topics: security
|
||||
- **declining_focus**: Topics fading from focus: gitops
|
||||
|
||||
## Signal Summary
|
||||
|
||||
| Metric | Baseline | Recent |
|
||||
|---|---|---|
|
||||
| Avg msg length | 1083 | 1531 |
|
||||
| Frustration rate | 4.0% | 6.0% |
|
||||
| Correction rate | 6.8% | 8.9% |
|
||||
| Action-first | 1.7% | 0.0% |
|
||||
| Top topics | ceph, proxmox, k8s | security, galera, ai_ml |
|
||||
| Avg msg length | 1192 | 2361 |
|
||||
| Frustration rate | 4.2% | 6.6% |
|
||||
| Correction rate | 7.1% | 8.8% |
|
||||
| Action-first | 1.5% | 0.5% |
|
||||
| Top topics | ceph, proxmox, k8s | security, backup, ceph |
|
||||
|
||||
## Recommendations
|
||||
|
||||
|
||||
@@ -3,24 +3,27 @@ title: Infrastruktur-Übersicht
|
||||
category: entities
|
||||
tags: [homelab, hardware, overview]
|
||||
created: "2026-04-28"
|
||||
modified: "2026-07-24"
|
||||
modified: "2026-09-26"
|
||||
---
|
||||
|
||||
# Infrastruktur-Übersicht
|
||||
|
||||
## Physikalische Hardware
|
||||
|
||||
### Proxmox Cluster (PVE 9.2.3)
|
||||
- **9 Nodes**, Quorum OK
|
||||
### Proxmox Cluster (PVE 9.2.20)
|
||||
- **8 Nodes**, Quorum OK (proxmox1+2 dauerhaft entfernt)
|
||||
- Hypervisoren in 10.0.20.x
|
||||
- Siehe [[systems/proxmox-cluster]]
|
||||
|
||||
### Ceph Cluster
|
||||
- **14 OSDs** (HDD + NVMe/SSD混合)
|
||||
- **17 OSDs** (10 HDD + 7 SSD), all up/in, v20.2.4
|
||||
- HEALTH_WARN (insecure key types — cosmetic)
|
||||
- 4 MONs (px5 leader, px4, a2-1, n5pro), MGR auf n5pro
|
||||
- Siehe [[systems/ceph-cluster]]
|
||||
|
||||
### RKE2 Kubernetes Cluster
|
||||
- **6 Nodes** (3 CP + 3 Worker), v1.35.6-rke2r1
|
||||
- worker-04 currently NotReady (VM 139 offline)
|
||||
- Alle Nodes schedulable (keine CP Taints)
|
||||
- Siehe [[systems/rke2-kubernetes]]
|
||||
|
||||
|
||||
@@ -22,8 +22,8 @@
|
||||
- [[entities/health-fitness]] — Gesundheits-Ziele, Ernährung
|
||||
|
||||
## Systems
|
||||
- [[systems/proxmox-cluster]] — PVE 9.2.3, 9 Nodes, Fluent Bit
|
||||
- [[systems/ceph-cluster]] — 14 OSDs, EC Pools, bekannte Probleme
|
||||
- [[systems/proxmox-cluster]] — PVE 9.2.20, 8 Nodes, Fluent Bit
|
||||
- [[systems/ceph-cluster]] — 17 OSDs, v20.2.4, 4 MONs, bekannte Probleme
|
||||
- [[systems/rke2-kubernetes]] — 6 Nodes, ArgoCD, CNPG, Workloads
|
||||
- [[systems/galera-maxscale]] — 3 Galera + 2 MaxScale, VIP .70
|
||||
- [[systems/loki-fluentbit]] — Logging Stack, 40+ Targets
|
||||
@@ -31,12 +31,19 @@
|
||||
- [[systems/hindsight]] — Semantic Memory, K8s, API
|
||||
- [[systems/monitoring]] — Prometheus, Grafana, HolmesGPT
|
||||
- [[systems/seafile]] — cloud.familie-schoen.com, Seafile 13.0.19
|
||||
- [[systems/laya]] — 421M Classifier CT152, choice/noul/scale, Session-Pre-Filter
|
||||
- [[systems/noris-ai]] — ai.noris.de LLMs, Embeddings, Image Gen
|
||||
- [[systems/frigate]] — NVR v0.18 CT151, Kameras, MQTT, HA Automation
|
||||
- [[systems/homeassistant]] — HAOS, MQTT, Automations, Notify
|
||||
- [[systems/paperless]] — Dokumentenmgmt, OIDC, OCR via noris AI
|
||||
- [[systems/sarah-hermes]] — VM107, zweites Hermes für Sarah (WebUI :8787, TG geplant)
|
||||
|
||||
## Concepts
|
||||
- [[concepts/network-architecture]] — 10.0.X.Y Schema, VLANs
|
||||
- [[concepts/gitops-workflow]] — IaC → Git → ArgoCD → Verify
|
||||
- [[concepts/credential-policy]] — 1Password, ESO, keine Secrets in Git
|
||||
- [[concepts/email-organization]] — Rechnungs-Organizer, Himalaya
|
||||
- [[concepts/memory-layer-architecture]] — 7-Schichten Memory Trennungsregeln
|
||||
|
||||
## Reference
|
||||
- [[reference/ip-map]] — IP → Host → Service Mapping
|
||||
@@ -54,6 +61,11 @@
|
||||
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
|
||||
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
|
||||
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
|
||||
- [[patterns/pve-oom-frozen-guest]] — Host-OOM friert Gast als Zombie ein, RSS-Kollaps-Signatur (PAT-010)
|
||||
- [[patterns/pve-notification-webhook]] — PVE Webhook→Telegram: Endpoint-Drift, Base64-Trap, Handlebars-Escape (PAT-011)
|
||||
- [[patterns/placement-guardrails]] — 70% RAM-Deckel + Swap-Verbot als Prometheus-Rules (PAT-012)
|
||||
- [[patterns/prometheus-config-rescue]] — grep -vn Gift/Crashloop, LXC-Rootfs-Direct-Mount-Rescue (PAT-013)
|
||||
- [[patterns/ceph-dead-nvme-resurrection]] — PCI-Reset→lvchange→chown-Falle, Weight-Management bei Reintegration (PAT-014)
|
||||
- [[patterns/skill-impact]] — Skill Modification Audit Trail
|
||||
|
||||
## Memory Layer Architektur
|
||||
|
||||
@@ -1,5 +1,134 @@
|
||||
# Memory Log
|
||||
|
||||
## [2026-09-30] k8s-cp01-relief-descheduler | Workload-Migration + Descheduler-Deployment
|
||||
- cp-01 bei 93% RAM → hindsight-api + hindsight-postgres + paperless nach worker-05 migriert (cordon/delete/uncordon). cp-01 RAM 93%→43%.
|
||||
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren bis konvergiert.
|
||||
- hindsight-api Live-Pin (Sep-17-Hotfix) war nie in Git → Live-Patch nötig.
|
||||
- Descheduler v0.30.0 deployt (raw Manifests, nicht Helm — Chart hat params/args-Bug + falsche API-Version). Policy: LowNodeUtilization (30/50/20 → 70/70/60), 2m Intervall, nodeFit auf Profile-Ebene.
|
||||
- RBAC-Lektion: Descheduler braucht `list/watch` auf `namespaces` in ClusterRole.
|
||||
- ArgoCD repoURL: `ssh://git@10.0.30.200:22/...` (mit `git@` User-Prefix, wie alle anderen Apps).
|
||||
- Solution Doc + INDEX + Wiki rke2-kubernetes.md aktualisiert. Commits b955d7e→8c1ecff.
|
||||
|
||||
## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur
|
||||
- Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen.
|
||||
- Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate.
|
||||
- SMART-Korrektur: Kingston SFYRDK2000G (osd.2-Träger) ist GESUND (wear 6 %, media_errors 0, spare 100 %, PoH 1469) — frühere "Austausch-Kandidatin"-These RETRAKIERT. Vorfall war Controller-/Fabric-Ebene. Neu beobachten: Temps 70/78 °C, T1-throttle 3×.
|
||||
- PAT-014 erweitert (Smart-Log-Abschnitt), ceph-cluster.md korrigiert, Commit cc2f237.
|
||||
|
||||
## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014)
|
||||
- osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin. *(Korrektur später am selben Tag: SMART-Audit → gesund, kein Austausch; siehe nächster Eintrag.)*
|
||||
- osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend.
|
||||
- ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben.
|
||||
- Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014.
|
||||
|
||||
## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013)
|
||||
- Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün.
|
||||
- Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok.
|
||||
- **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate.
|
||||
- Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit <SHA>.
|
||||
|
||||
## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012)
|
||||
- Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok.
|
||||
- Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom).
|
||||
- Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172.
|
||||
|
||||
## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011)
|
||||
- Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON).
|
||||
- Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200.
|
||||
- Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body.
|
||||
- Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales).
|
||||
|
||||
## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen
|
||||
- proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%.
|
||||
- Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets.
|
||||
- VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP.
|
||||
- Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt.
|
||||
|
||||
## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet
|
||||
- GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md.
|
||||
- Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc.
|
||||
- proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen.
|
||||
- Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert.
|
||||
|
||||
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
|
||||
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
|
||||
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
|
||||
- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX
|
||||
- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared
|
||||
- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr
|
||||
- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return
|
||||
- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010)
|
||||
|
||||
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
|
||||
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
|
||||
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
|
||||
- E2E verifiziert: Login 200, Agent-Antwort via vllm/release/glm-5-2 @ ai.noris.de
|
||||
- PITFALL: hermes-agent pyproject pinnt Core-Deps hinter python_version>='3.14'-Markers → auf Python 3.12 installiert `pip install -e .` NULL Deps ("AIAgent not available"). Fix: manuelle Pin-Installation + Missing-Import-Loop bis `import run_agent`
|
||||
- Backup: täglich 23:00 Job backup-2e8a34e3-66cb (all=1) deckt VM107; manueller Verify TASK OK 48s
|
||||
- Wiki: systems/sarah-hermes.md, ip-map.md erweitert; iac-homelab commits 0d5f017/1ce4816
|
||||
- OFFEN: Telegram-Bot blockiert auf BotFather-Token (Sarah/Dominik)
|
||||
|
||||
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
|
||||
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
|
||||
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
|
||||
- Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min)
|
||||
- Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05
|
||||
- Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2)
|
||||
- Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes)
|
||||
- Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md
|
||||
|
||||
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
|
||||
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
|
||||
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
|
||||
- Gitea Pod + schoenkitchen runner durch Pod-Delete wiederhergestellt
|
||||
- ArgoCD Audit: 17/25 Apps via Gitea SSH, 20/25 Synced+Healthy
|
||||
- Known Issues: `gitea`/`authelia` Health=Progressing (Ingress-LB-IP Gap), `immich` doppelt verwaltet, `gitea-config` Deployment Drift
|
||||
- `DEFAULT_ACTIONS_URL=https://gitea.com` deprecated → Helm override auf `self` empfohlen
|
||||
- Wiki `systems/gitea.md` aktualisiert: Runner Status, Incident, ArgoCD Issues
|
||||
- Solution Doc: `docs/solutions/bug-fixes/2026-09-27-gitea-csi-rbac-volumeattachment-fix.md`
|
||||
|
||||
## [2026-09-27] architecture | Memory Layer Restructuring + Laya Session-Nutzung
|
||||
- USER.md bereinigt: Infra-Fakten entfernt, nur noch User-Preferences/Safety-Rules (1.087/1.375 chars)
|
||||
- MEMORY.md ausgedünnt: 2.145→1.664 chars, alle Infra-Details zeigen auf Wiki-Seiten mit `→Wiki` Pointern
|
||||
- 4 neue Wiki-Seiten: systems/noris-ai, systems/frigate, systems/homeassistant, systems/paperless
|
||||
- Neues Concept: concepts/memory-layer-architecture — Entscheidungsbaum "was wohin gehört"
|
||||
- Hindsight Audit: massiv überladen mit veralteter Infra-Topologie. Going-Forward-Policy: nur noch semantische Pointer + Entscheidungen
|
||||
- LCM gesund: 7.590 messages, 34 DAG nodes, 21.2:1 compression ratio
|
||||
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
|
||||
|
||||
## [2026-09-27] architecture | Laya Email-Organizer Migration + Session-Nutzung
|
||||
- Rechnungen-Organizer Cron `f773f8c23230` von LLM-Agent → `no_agent` Script mit Laya migriert
|
||||
- 12 Kategorien, 3-Schichten-Safety (Confidence-Gate + Subject-Validierung + Move-Erfolg)
|
||||
- Globale Ordner (Rechnungen/Bestellungen/Gutschriften/Gutscheine) statt Monatssortierung
|
||||
- Wiki-Seite `systems/laya.md` erstellt mit Usage Guide für Session-Nutzung
|
||||
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
|
||||
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
|
||||
|
||||
## [2026-09-17] workflow | CT108-Endausbau: Census, Ghost-Router CT99999, Alias-Repair, Stop
|
||||
- **Census (auth):** K8s=25 / CT108=31 / gemeinsam=23 — 22 Tips identisch, dominik/memory=K8s-Superset (enthält CT-Tip de97cd6d), nur-CT=8× PoC-Müll. Archiv: 29 Bare-Bundles 143 MB unter `/home/debian/git-archive/ct108-final/` (2 Failures = legitim leere Repos).
|
||||
- **Ghost-Router enttarnt:** CT99999 (Traefik-LXC auf proxmox7, 10.0.60.10) terminiert TLS für *.familie-schoen.com und forwardet plain HTTP an .203. `git.familie-schoen.com` → .105:3000 WAR der letzte Live-Konsument von CT108; `git.schoen.codes` lief längst per Double-Hop aufs K8s. Fix: gitea-service-Upstream → .203 (Backup `explicit-http.yml.bak-hermes-20260917`) + Alias-Host im K8s-Ingress (Commit `9e9b6ee`). Beide Hostnamen jetzt v1.27.0.
|
||||
- **Stop vollzogen (genehmigt):** `pct stop 108` auf ms-a2-2 (10.0.20.93) — connection-refused-Beweis, Fleet 22/25 Synced, Runner unversehrt. Wiki-Lügen korrigiert (ip-map behauptete „stopped" seit Wochen).
|
||||
- **Lessons:** 1P-SA braucht je Call `--vault`; `op read --reveal` existiert nicht (stdout=Secret); `op item get --reveal` maskiert nur Display (JSON-Captures intakt); `/repos/search` = `{ok,data}`-Envelope; blankes Token erzeugt glaubwürdig LEERE Census (Fast-Fehlentscheidung „K8s hat nur 2 Repos").
|
||||
- Docs: `docs/solutions/workflows/2026-09-17-ct108-full-decommission-census-router-topology.md` · Wiki: systems/gitea.md, reference/ip-map.md
|
||||
- Offen (je Freigabe): Zombie-Secret `argocd-repo-credentials` löschen; immich-Zwillings-App (toter rendered-Dump vom 01.08., Live gehört immich-config) entfernen.
|
||||
|
||||
## [2026-09-17] fix | Merge-Day-Kampagne abgeschlossen: Konvergenz + Autosync-Rennen + Live-State-Arbitrage
|
||||
- **Konvergenz DONE:** Merge `d76d3e8` (fork 46fa168, 24.07.) + Fixups `4987341`/`8c741c5` auf BEIDEN Remotes (CT108 + K8s-Gitea SSH 10.0.30.200). Ahead37/behind63-Narrativ endgültig begraben (Cache-Phantom). Single-Remote-Ziel erreicht: beide Tipps identisch.
|
||||
- **Autosync-Renne verarbeitet (4 Minentypen):** authelia Duplikat-Volume (union-merge) entfernt; homepage `authelia-auth`-Middleware restauriert (Jul-Entscheidung ging nie live); paperless plaintext-OIDC-Secret GELÖSCHT statt Wert-Rollback (ESO-Ownership seit 24.08., Quelle 1P); newborn-App ceph-csi-cephfs eingefroren (autosync-Block aus Git entfernt — Helm-Release 3.17.0 wartet auf Adoption; rbd-chart 3.10.1 weiterhin ungoverned, Backlog).
|
||||
- **3 Sync-Failures via Live-State-Arbitrage geheilt (alle Synced/Healthy @ 8c741c5):** (1) backups: VSC-CRD-Flavor-Falle — RKE2-Addon-CRD ist FLAT (driver/deletionPolicy top-level, `spec:` verboten!), Jul-Files waren für Upstream-Flavor korrekt; beide Files geflattet, cephfs-snapclass erstmals LIVE ERZEUGT. (2) databases: CNPG grow-only — Git auf 100Gi/20Gi hochaligned (Shrink verboten). (3) hindsight: VCT storageClassName immutable (hdd→flash aligniert), Live-Hotfixes gespiegelt (Node-Pin cp-01, ReadinessProbe draußen, secretKeyRef statt Klartext-PW). hindsight-postgres-0 rotierte sauber, PVC 47d Bound, API pollt.
|
||||
- **Push-Learned:** canonical-Push braucht EXPLIZITEN Key (`GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o IdentitiesOnly=yes"`) — kein ~/.ssh/config vorhanden, Default-Key lehnt ab.
|
||||
- **Chronisch (prä-merge, offen):** kube-prometheus-stack Synced/FAILED (CRD-Annotation >256kB); residual OutOfSync gitea-config (Deployment/gitea-runner) + immich-config (SA/CM-Reste) = Alt-Backlog, keine Regression.
|
||||
- **Finale Ordnung (nächste Schritte):** ① ArgoCD-Sources auf ssh://git@10.0.30.200:22 flippen (incl. Secret argocd-repo-credentials) ② CT108 final stoppen (explizite Freigabe Dominik) ③ CSI-Adoption ④ kpstack-CRD-Fix.
|
||||
- Docs: `docs/solutions/workflows/2026-09-17-argocd-autosync-race-four-mine-types.md` + `docs/solutions/bug-fixes/2026-09-17-live-state-arbitration-crd-flavors-hotfix-mirroring.md` · Morgen-Doc (.202-Empfehlung) korrigiert auf .200.
|
||||
|
||||
## [2026-09-17] bug-fix | Ghost-Instance-Regression: CT108-Gitea als ArgoCD-Source + gitea-backup RCA-Quality
|
||||
- gitea-backup-29826930 (03:30Z) failed: BackoffLimitExceeded, 3 Instant-Crashes in 73s. Zwei Auto-RCAs attribuierten auf 1P/ESO-Rate-Limit — STRUKTURELL widerlegt (crashing init-container gitea-files konsumiert keine Secrets; mounted Secrets intakt).
|
||||
- Echter Fund: ArgoCD-Apps (root/proxy/paperless/schoenkitchen/gitea-config) zogen von http://10.0.30.105:3000 = CT108-Ghost (nach Dekommissionierung 24.07. wieder eingeschaltet, stale Mirror ohne Hardening-Commit 2014673) → "Synced" maskierte fehlende failedJobsHistoryLimit-Felder live.
|
||||
- Sofortmassnahme: manueller Rerun gitea-backup-manual-161538 SUCCESS (201,7 MB, S3-Upload verifiziert), Failed-Job gelöscht, Alert clear.
|
||||
- Offen (awaiting owner): ArgoCD-Sources auf ssh://git@10.0.30.202:22 flippen, Repo-Divergenz ahead37/behind63 + tote HTTPS-Auth forensisch, CT108 final stoppen, ESO-Nachtsättigung (00:00Z-Fenster, cf. PAT-003) analysieren.
|
||||
- Doc: `docs/solutions/bug-fixes/2026-09-17-gitops-ghost-instance-regression-masked-hardening.md` · Wiki: systems/gitea (CT108-Status), concepts/gitops-workflow (Single-Remote-Regel)
|
||||
- **Abend-Phase:** SSH-Deny-Rootcause = keine Keys/Tokens im K8s-Gitea-DB registriert (Opfer der 01.08.-Migration) → via 1P-Token (hermes-gitops) re-registriert: User-Key + ArgoCD-Deploy-Key (beide end-to-end verifiziert, push dry-run OK). Echter Branch-Split am 24.07. entdeckt: K8s-Gitea-main=29.07.-Stand (b9d7441, 70 Commits incl. mariadb:11.4-Fix für gitea-backup), CT108/local=38 Commits ab 04.09. Hardening TTL=86400 + Exit-42-Guard via ArgoCD gelanded (a0c9b6b). Heutiger DB-Dump validiert (116 Tables, kompletter Trailer). Nächste Schritte: Historien-Konvergenz → Source-Flip → CT108-Stop (Freigabe).
|
||||
|
||||
## [2026-08-30] retro | Compound Learning Retrospective (Last 30 Days)
|
||||
- Reviewed sessions from Jul 31 – Aug 30, 2026
|
||||
- **4 new solution docs** written by subagents:
|
||||
@@ -161,3 +290,11 @@
|
||||
- Created 2 concept pages: gitops-workflow, credential-policy
|
||||
- Rewrote index.md with new structure + Memory Layer Architecture table
|
||||
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
|
||||
|
||||
## [2026-09-29] update | VM302 Zombie-Recovery nach Migration
|
||||
- VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot
|
||||
- Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause
|
||||
- Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min
|
||||
- Endstand: 3/3 Synced, Primary, MaxScale alle Server Running
|
||||
- Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md
|
||||
- Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# PAT-014 — Dead-NVMe-Resurrection für Ceph-OSD
|
||||
|
||||
## Trigger
|
||||
Ceph-OSD auf externem/non-Corosync-Host startet nicht: KernelDevice errno-5 I/O-Errors beim
|
||||
Label-Lesen, Symlink `/var/lib/ceph/osd/ceph-*/block` ins Leere, VG „verschwindet".
|
||||
|
||||
## Root Cause
|
||||
NVMe-Controller im PCIe-Fabric gestorben (state=dead, Namespace 0B) — NICHT das Medium selbst.
|
||||
Nach PCI-reset kehrt der Controller oft unter NEUER Nummer zurück (nvme0 → nvme2!), alte
|
||||
Device-Mapper-Tabelle bleibt stale.
|
||||
|
||||
## Resurrection-Sequence (verifizierte Reihenfolge)
|
||||
1. **Korrekte PCI-Addr finden**: über sysfs-Pfad des Namespaces (`/sys/class/block/nvmeXnY/device`),
|
||||
NICHT raten — erster Versuch traf den falschen (gesunden!) Controller.
|
||||
2. `echo 1 > /sys/bus/pci/devices/<ADDR>/remove && echo 1 > /sys/bus/pci/rescan`
|
||||
→ Controller kommt als neue Instanz zurück, Namespaces wieder da.
|
||||
3. `pvscan --cache` → VG/LV wieder sichtbar.
|
||||
4. **Stale DM-Table**: `lvchange -an <vg>/<lv> && lvchange -ay -K <vg>/<lv>`
|
||||
5. **Lesetest**: `dd if=/dev/<vg>/<lv> bs=4M count=8 of=/dev/null` — bei weiteren errno-5:
|
||||
NAND/Media defekt → OSD out lassen, Drive tauschen.
|
||||
6. **Permissions-Falle nach lvchange**: udev setzt Owner root →
|
||||
`chown ceph:ceph /dev/mapper/<dm-name>; chmod 660` (sonst `bdev open: (13) Permission denied`).
|
||||
7. `systemctl reset-failed ceph-osd@<id> && systemctl start ceph-osd@<id>`
|
||||
8. `ceph osd in osd.<id>` + Weight restaurieren.
|
||||
|
||||
## Begleitregeln
|
||||
- Klein/niedergewichtetes OSD niemals mit Weight 1.0 reintegrieren, solange Pools an
|
||||
Full-Ratios kratzen → sonst `backfillfull`/`backfill_toofull`-Blockade (Fall osd.3:
|
||||
re-in@1.0 → 96 % voll instant → 13 PGs blocked; Fix: `crush reweight osd.3 0.05`).
|
||||
- Externe Non-Corosync-Hosts in `/etc/hosts` ALLER PVE-Nodes pflegen (GUI-500
|
||||
„hostname lookup failed"), oder echter DNS-Record.
|
||||
|
||||
## Verified
|
||||
2026-09-30: osd.2 revived (Kingston SFYRDK2000G, PCI 03:00.0, → nvme2), 495 GiB, Weight 1.0;
|
||||
Cluster 17/17 up/in; Degraded 0,78 % fallend. osd.3 stabilized @ weight 0.05.
|
||||
|
||||
## Smart-Log-Abgleich 2026-09-30 (korrigiert frühere Alters-These)
|
||||
Kingston SFYRDK2000G (nvme0n2, trägt osd.2): percentage_used **6 %**, media_errors **0**,
|
||||
available_spare 100 %, unsafe_shutdowns 2, power_on_hours 1469, power_cycles 3.
|
||||
Geschwisterplatte nvme1n1 identisches Profil (6 % wear, 0 media errors).
|
||||
⇒ Drive ist MEDIZINISCH GESUND — der Vorfall war rein Controller-(PCI)-Ebene, kein
|
||||
Media-Verschleiß. **Kein Austausch nötig.** Einziges Beobachtungsfeld: Temperatur 70 °C /
|
||||
Sensor2 78 °C (thermisches Throttling T1 3× aktiviert) — Kühlung prüfen wäre sinnvoll,
|
||||
aber keine Akutgefahr.
|
||||
@@ -0,0 +1,23 @@
|
||||
# PAT-012 — Placement Guardrails (RAM-Deckel + Swap-Verbot)
|
||||
|
||||
**Trigger**: Neue Gastplatzierung oder Kapazitätsfrage auf einem PVE-Node.
|
||||
|
||||
**Policy**:
|
||||
- Committed RAM pro Hypervisor ≤ 70 % physikalisch. Über Deckel = "voll",
|
||||
Guests redistribuieren BEVOR neue Last kommt.
|
||||
- Resident Swap > 100 MiB für >10 min = Politikverstoß → Placements prüfen,
|
||||
nie als Normalzustand akzeptieren.
|
||||
|
||||
**Enforcement (live)**: Gruppe `placement_guardrails` in CT141
|
||||
(`/opt/monitoring/prometheus/rules/alerting_rules.yml`):
|
||||
- `PVEPlacementCeilingBreached` (ratio > 0.70, for 15m, warning)
|
||||
- `HVResidentSwapNonzero` (SwapTotal-Free > 100MiB, for 10m, warning)
|
||||
|
||||
**Vor jeder Platzierung**: Ratio-Query ziehen
|
||||
(`pve_memory_usage_bytes{id=~"node/.*"} / pve_memory_size_bytes{id=~"node/.*"}`)
|
||||
— >70 %-Nodes meiden, Headroom liegt primär auf ms-a2-1/ms-a2-2.
|
||||
|
||||
Baseline-Rollout 30.09.: proxmox3 80,6 / p4 79,3 / p5 75,5 / p6 72,6 %
|
||||
über Deckel (Warnburst erwartet = Rebalancing-Backlog, kein Emergency).
|
||||
|
||||
Ref: docs/solutions/architecture/2026-09-30-placement-guardrails-ram-ceiling.md
|
||||
@@ -0,0 +1,28 @@
|
||||
# PAT-013 — Prometheus Config Rescue (Crashloop + LXC Rootfs Channel)
|
||||
|
||||
**Trigger**: Prometheus crashloop nach Config-Rewrite, oder pct exec unbrauchbar.
|
||||
|
||||
**Iron Rules**:
|
||||
1. NIEMALS `grep -n`-Ausgabe als Rewrite-Source — Nummernpräfixe verseuchen die Datei.
|
||||
Nutze `grep -v` (ohne -n) oder Python/awk.
|
||||
2. YAML validieren BEVOR Container recreate (sonst Crashloop = totale Blindheit).
|
||||
3. Beim Löschen von Targets/IPs ALLE Locations sweeppen — Dubletten leben gern in
|
||||
mehreren Jobs derselben Datei (node_exporter-Liste ≠ blackbox-Liste).
|
||||
|
||||
**Rescue-Kanal (universell für LXC auf RBD)**:
|
||||
```bash
|
||||
# Auf dem Hyper visor, der den CT hostet:
|
||||
lsblk # rbdX finden (size matchen)
|
||||
mount /dev/rbdX /mnt/rescue
|
||||
# Dateien direkt editieren unter /mnt/rescue/opt/...
|
||||
umount /mnt/refuge
|
||||
# im CT: docker compose up -d --force-recreate <svc>
|
||||
```
|
||||
|
||||
**Symptom-Signature Crashloop**: `docker logs prometheus` → "yaml: line N: mapping
|
||||
values are not allowed in this context" = klassisches NN:-Präfix-Gift.
|
||||
|
||||
**Stale-Series-Note**: gelöschte Targets erscheinen sekundenweise weiter in Queries —
|
||||
erst nach Scrape-Zyklus-Gap re-checken, dann Erfolg erklären.
|
||||
|
||||
Ref: docs/solutions/bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md
|
||||
@@ -0,0 +1,22 @@
|
||||
# PAT-011 — PVE Webhook Notifications Pipeline Repair
|
||||
|
||||
**Trigger**: PVE notifications (esp. vzdump) reach Telegram not / test returns 500.
|
||||
|
||||
**Signature diagnosis ladder**:
|
||||
1. `pvesh create /cluster/notifications/targets/<name>/test` — error taxonomy:
|
||||
- `Connection refused` → transport/down (check URL target alive)
|
||||
- `failed to render webhook body` → template broken (syntax/base64)
|
||||
- `http status: 400` → bridge rejected payload (escaping!)
|
||||
2. Listener-Sweep über VLAN: `for ip in .xx…; do /dev/tcp/$ip/port probe; done`
|
||||
3. Byte-Level-Truth via temp mini-sniffer (redirect URL, capture, restore!).
|
||||
|
||||
**Hard rules**:
|
||||
- Mutations an `/etc/pve/notifications.cfg` NUR via `pvesh set` (hand-edits poison
|
||||
global deserialization!). Vorher Snapshot.
|
||||
- `body`/header-values/secrets = **base64 blobs**: `printf '%s' tpl | base64 -w0`.
|
||||
- Handlebars helpers: `{{escape title}}` (Space!), NICHT `{{escape:title}}`.
|
||||
- Immer `{{escape title}}`/`{{escape message}}` verwenden — rohe Interpolation bricht
|
||||
bei Apostrophen/Multiline (Backup-Reports!).
|
||||
- Sniffer-Redirect URL IMMER restaurieren.
|
||||
|
||||
Reference: `docs/solutions/bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md`
|
||||
@@ -0,0 +1,49 @@
|
||||
---
|
||||
pattern_id: PAT-010
|
||||
title: "PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature)"
|
||||
category: infrastructure
|
||||
severity: high
|
||||
status: active
|
||||
first_observed: 2026-09
|
||||
last_updated: 2026-09-30
|
||||
related_systems: [proxmox-cluster, rke2-kubernetes]
|
||||
related_solution_docs: [docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md]
|
||||
related_skills: [rke2-cluster-administration, systematic-debugging]
|
||||
---
|
||||
|
||||
# PAT-010: PVE host OOM-kill freezes guest VM as zombie
|
||||
|
||||
## Symptom
|
||||
- K8s node NotReady, kubelet heartbeat stops abruptly
|
||||
- PVE shows VM `running`, but: no ping, no SSH, qemu-guest-agent dead,
|
||||
tap-interface TX counters static
|
||||
- **Signature:** `qm status <vmid> --verbose` → kvm RSS collapses to a
|
||||
tiny fraction (<5%) of assigned RAM — guest kernel no longer touches
|
||||
its memory
|
||||
- Downstream alerts (DS misscheduled, workload CrashLoops) fire, but
|
||||
NOTHING alarms on the actual OOM
|
||||
|
||||
## Root Cause
|
||||
Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks
|
||||
the largest anon-RSS process = biggest kvm. After TWO consecutive kills of
|
||||
the same guest (HA auto-restarts in between), the revived QEMU comes up
|
||||
but the guest kernel stays inert → zombie VM.
|
||||
|
||||
## Mitigation
|
||||
1. Relieve host pressure FIRST (move movable tenants away) — otherwise
|
||||
the unfreeze re-boots the guest into the same thrash.
|
||||
2. Hard cycle the VM: `qm stop` + `qm start` (respect HA guards).
|
||||
3. Re-query `ha-manager status` afterwards — HA may relocate the VM.
|
||||
4. Uncordon the K8s node; orphan DS pods self-reconcile.
|
||||
|
||||
## Prevention
|
||||
- Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/
|
||||
IaC; swap is a buffer, not capacity.
|
||||
- PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently
|
||||
invisible to alerting (noticed only via downstream K8s symptoms).
|
||||
|
||||
## Evidence
|
||||
- 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed
|
||||
rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second
|
||||
revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration +
|
||||
hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.
|
||||
+21
-21
@@ -3,54 +3,54 @@ title: IP-Map (Quick Reference)
|
||||
category: reference
|
||||
tags: [ip, network, reference, quick-lookup]
|
||||
created: "2026-07-24"
|
||||
modified: "2026-07-24"
|
||||
modified: "2026-09-30"
|
||||
---
|
||||
|
||||
# IP-Map
|
||||
|
||||
## Proxmox Hosts (10.0.20.x) — 9 Nodes
|
||||
## Proxmox Hosts (10.0.20.x) — 8 Nodes
|
||||
| IP | Hostname | Node ID | Notes |
|
||||
|----|----------|---------|-------|
|
||||
| 10.0.20.20 | proxmox2 | 3 | |
|
||||
| 10.0.20.30 | proxmox3 | 4 | |
|
||||
| 10.0.20.40 | proxmox4 | 2 | |
|
||||
| 10.0.20.50 | proxmox5 | 5 | MON, MGR |
|
||||
| 10.0.20.40 | proxmox4 | 2 | MON |
|
||||
| 10.0.20.50 | proxmox5 | 5 | MON (leader) |
|
||||
| 10.0.20.60 | proxmox6 | 6 | |
|
||||
| 10.0.20.70 | proxmox7 | 7 | MON, Traefik CT99999 |
|
||||
| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002 |
|
||||
| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0 |
|
||||
| 10.0.20.91 | n5pro | 8 | Templates 9000/9001/9002, MON, MGR (active) |
|
||||
| 10.0.20.92 | ms-a2-1 | 9 | OSD 13 (SSD 1.8TB), GPU 1002:13c0, MON |
|
||||
| 10.0.20.93 | ms-a2-2 | 1 | OSD 14 (SSD 1.8TB), GPU 1002:13c0 |
|
||||
|
||||
> ⚠️ n5pro (10.0.20.91, nodeid 8) ≠ proxmox7 (10.0.20.70, nodeid 7) — separate Nodes!
|
||||
> proxmox1 (10.0.20.10, nodeid 1) wurde dauerhaft entfernt (2026-07-24).
|
||||
> proxmox1 (10.0.20.10) und proxmox2 (10.0.20.20) wurden dauerhaft entfernt.
|
||||
> Immer `pvecm nodes` für kanonische Liste prüfen.
|
||||
|
||||
## Kubernetes Nodes (10.0.30.5x-6x)
|
||||
| IP | Node | Role |
|
||||
|----|------|------|
|
||||
| 10.0.30.51 | cp-01 | Control Plane |
|
||||
| 10.0.30.52 | cp-02 | Control Plane |
|
||||
| 10.0.30.53 | cp-03 | Control Plane |
|
||||
| 10.0.30.63 | worker-01 | Worker |
|
||||
| 10.0.30.64 | worker-04 | Worker (GPU ✅) |
|
||||
| 10.0.30.65 | worker-05 | Worker (GPU defekt) |
|
||||
| IP | Node | Role | Status |
|
||||
|----|------|------|--------|
|
||||
| 10.0.30.51 | cp-01 | Control Plane | Ready |
|
||||
| 10.0.30.52 | cp-02 | Control Plane | Ready |
|
||||
| 10.0.30.53 | cp-03 | Control Plane | Ready |
|
||||
| 10.0.30.61 | worker-01 | Worker | Ready |
|
||||
| 10.0.30.64 | worker-04 | Worker (GPU) | **NotReady** |
|
||||
| 10.0.30.65 | worker-05 | Worker (GPU) | Ready |
|
||||
|
||||
## Database Layer (10.0.30.7x-8x)
|
||||
| IP | Host | Service |
|
||||
|----|------|---------|
|
||||
| 10.0.30.70 | — | MaxScale VIP (keepalived) :3306 |
|
||||
| 10.0.30.71 | VM300 | Galera db1 (ms-a2-1) |
|
||||
| 10.0.30.72 | VM301 | Galera db2 (proxmox3) |
|
||||
| 10.0.30.73 | VM302 | Galera db3 (proxmox6) |
|
||||
| 10.0.30.71 | VM300 | Galera db1 (n5pro) |
|
||||
| 10.0.30.72 | VM301 | Galera db2 (proxmox6) |
|
||||
| 10.0.30.73 | VM302 | Galera db3 (ms-a2-2) |
|
||||
| 10.0.30.81 | VM310 | MaxScale-01 Admin :8989 |
|
||||
| 10.0.30.82 | VM311 | MaxScale-02 Standby |
|
||||
|
||||
## Infrastructure VMs (10.0.30.x)
|
||||
| IP | Host | Service |
|
||||
|----|------|---------|
|
||||
| 10.0.30.66 | sarah-hermes (VM107, n5pro) | Sarahs Hermes: WebUI :8787 (LAN-only, pw) + TG-Bot (geplant) |
|
||||
| 10.0.30.99 | CT111 | Immich (n5pro), Migration zu K8s geplant |
|
||||
| 10.0.30.100 | ubuntu | Physischer Ubuntu-Node (Ceph OSDs, ZFS pool01_n2_redundant) |
|
||||
| 10.0.30.105 | — | ~~Gitea~~ DECOMMISSIONED (CT108, stopped) |
|
||||
| 10.0.30.105 | — | ~~Gitea~~ CT108 GESTOPPT 2026-09-17 (Archive: /home/debian/git-archive/ct108-final/) |
|
||||
| 10.0.30.124 | VM200 | IaC Runner (DHCP), SSH `debian` |
|
||||
| 10.0.30.141 | CT141 | Monitoring (Prometheus/Grafana) |
|
||||
| 10.0.30.145 | CT145 | Voice Pipeline (PJSIP :5060) |
|
||||
@@ -76,7 +76,7 @@ modified: "2026-07-24"
|
||||
## DMZ / Reverse Proxies (10.0.60.x)
|
||||
| IP | Host | Service |
|
||||
|----|------|---------|
|
||||
| 10.0.60.10 | CT9999 | Traefik Reverse Proxy (root/[REDACTED]) |
|
||||
| 10.0.60.10 | CT99999 | Traefik Outer-Proxy (root/[REDACTED]); Terminiert *.familie-schoen.com + leitet schoen.codes als Plain-HTTP an 10.0.30.203 |
|
||||
|
||||
## Related
|
||||
- [[concepts/network-architecture]]
|
||||
|
||||
+47
-52
@@ -3,37 +3,44 @@ title: Ceph Cluster
|
||||
category: systems
|
||||
tags: [ceph, storage, rbd, ec-pool, osd]
|
||||
created: "2026-07-24"
|
||||
modified: "2026-07-25"
|
||||
modified: "2026-09-26"
|
||||
---
|
||||
|
||||
# Ceph Cluster
|
||||
|
||||
## Overview
|
||||
- **Cluster ID**: 204c8171-e0b1-4f40-9de2-a7cfe4ef68d9
|
||||
- **Health**: HEALTH_OK (recovery complete after OSD 0+2 drain, 0.3% misplaced settling)
|
||||
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH 2026-07-25), 3 MONs (proxmox5/7/4), MGR on proxmox5
|
||||
- **OSDs**: 13 (8 SSD, 5 HDD), all up/in — OSDs 0+2 destroyed+purged 2026-07-25
|
||||
- **Capacity**: ~22 TiB total, 6.0 TiB used
|
||||
- **Health**: HEALTH_WARN — "Monitors are configured to allow creation of insecure key types" (cosmetic, CVE-2025-30156 fixed)
|
||||
- **Version**: 20.2.4 (tentacle) — all 17 OSDs
|
||||
- **Nodes**: 8 Proxmox hosts (proxmox2 removed from CRUSH), 4 MONs (proxmox5, proxmox4, ms-a2-1, n5pro), MGR on n5pro (standbys: px5/6/7/a2-1)
|
||||
- **OSDs**: 17 (10 HDD, 7 SSD), all up/in
|
||||
- **Capacity**: ~33 TiB total, 9.0 TiB used, 24 TiB avail
|
||||
- **Pools**: 13 pools, 533 PGs (532 active+clean, 1 scrubbing)
|
||||
|
||||
## OSD Layout
|
||||
|
||||
| OSD | Class | Size | Host | Reweight | Notes |
|
||||
|-----|-------|------|------|----------|-------|
|
||||
| 0 | ssd | 188 GB | proxmox2 | — | **DESTROYED 2026-07-25** (92% wear) |
|
||||
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | Large HDD |
|
||||
| 2 | ssd | 233 GB | proxmox2 | — | **DESTROYED 2026-07-25** (slow ops, 81% full) |
|
||||
| 3 | ssd | 238 GB | proxmox4 | 0.95 | |
|
||||
| 4 | ssd | 233 GB | proxmox3 | 0.95 | |
|
||||
| 5 | ssd | 238 GB | proxmox5 | 0.90 | 80% full |
|
||||
| 6 | hdd | 2.8 TiB | ubuntu | 1.0 | Large HDD |
|
||||
| 7 | hdd | 500 GB | proxmox7 | 1.0 | Was 0.80, reweighted 2026-07-24 |
|
||||
| 8 | hdd | 2.8 TiB | ubuntu | 1.0 | BlueFS spillover |
|
||||
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
|
||||
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
|
||||
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
|
||||
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
|
||||
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
|
||||
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
|
||||
| 7 | hdd | 931 GB | proxmox7 | 0.95 | |
|
||||
| 8 | hdd | 3.6 TiB | ubuntu | 1.0 | |
|
||||
| 9 | ssd | 1.9 TiB | n5pro | 1.0 | |
|
||||
| 10 | hdd | 300 GB | proxmox6 | 1.0 | Very small HDD |
|
||||
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | Large HDD |
|
||||
| 10 | hdd | 931 GB | proxmox6 | 0.95 | |
|
||||
| 11 | hdd | 2.8 TiB | n5pro | 1.0 | |
|
||||
| 12 | ssd | 1.9 TiB | n5pro | 1.0 | |
|
||||
| 13 | ssd | 1.8 TiB | ms-a2-1 | 1.0 | |
|
||||
| 14 | ssd | 1.8 TiB | ms-a2-2 | 1.0 | New 2026-07-24, nvme0n1 |
|
||||
| 13 | ssd | 1.8 TiB | ms-a2-1 | 0.95 | |
|
||||
| 14 | ssd | 1.8 TiB | ms-a2-2 | 0.95 | |
|
||||
| 15 | ssd | 1.8 TiB | ubuntu | 1.0 | New |
|
||||
| 17 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
|
||||
| 18 | hdd | 3.6 TiB | ubuntu | 1.0 | New |
|
||||
|
||||
> OSDs 0+2 (old proxmox2) destroyed 2026-07-25. osd.2 reassigned to ubuntu host as new SSD.
|
||||
> OSDs 15, 17, 18 added since last wiki update (ubuntu host expanded).
|
||||
|
||||
## Pools
|
||||
|
||||
@@ -44,7 +51,7 @@ modified: "2026-07-25"
|
||||
| 3 | vm_disks | replicated | 3 | 2 | 2 (ssd) | 128 | autoscale on |
|
||||
| 4 | .mgr | replicated | 3 | 2 | 2 (ssd) | 1 | |
|
||||
| 5 | rbd | replicated | 3 | 2 | 1 (hdd) | 32 | autoscale on |
|
||||
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true (was 120, equalized to 112) |
|
||||
| 6 | hdd_disk | replicated | 3 | 2 | 1 (hdd) | 112 | nopgchange=true |
|
||||
| 7 | tm_disks | replicated | 2 | 2 | 1 (hdd) | 128 | target_size 2TiB |
|
||||
| 8 | media_ec | erasure 4+1 | 5 | 4 | 3 (hdd, osd-level) | 128 | ec_overwrites |
|
||||
| 9 | media_meta | replicated | 3 | 2 | 0 (any) | 32 | |
|
||||
@@ -58,46 +65,34 @@ modified: "2026-07-25"
|
||||
|
||||
## Known Issues
|
||||
|
||||
### Weight Imbalance Causing Placement Failures (2026-07-24)
|
||||
HDD hosts have extreme weight disparity: n5pro=10.15TB, ubuntu=5.49TB, proxmox7=0.50TB, proxmox6=0.30TB.
|
||||
CRUSH host-level selection (rule 1) often picks only 2 of 4 HDD hosts → up sets with 2 OSDs instead of 3.
|
||||
Result: PGs stuck in `active+clean+remapped` because up set < min_size.
|
||||
### HEALTH_WARN: Insecure Key Types (2026-09-26)
|
||||
Monitors allow insecure key types. Cosmetic warning — CVE-2025-30156 already fixed in 20.2.4.
|
||||
Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not already set).
|
||||
|
||||
**Mitigation (2026-07-24)**:
|
||||
1. Reweighted osd.7 from 0.80 → 1.0 → fixed EC pool 8.3d (NONE → osd.7)
|
||||
2. Equalized pool 6 pg_num 120 → 112 + nopgchange=true
|
||||
3. Manual pg-upmap for stuck PGs: 5.13 → [1,6,7], 6.6c → [11,8,7], 6.58 → [11,6,7]
|
||||
4. All `clean+remapped` eliminated. Triggered rebalancing wave (43 PGs backfilling at 26 MiB/s).
|
||||
### Small SSDs causing reweightdown
|
||||
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
|
||||
Consider removing from CRUSH or replacing with larger drives.
|
||||
|
||||
**Long-term**: Small HDDs (osd.7 0.5TB, osd.10 0.3TB) cause CRUSH placement failures. Replace with larger disks or create separate CRUSH root for large HDDs only.
|
||||
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
|
||||
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
|
||||
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
|
||||
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
|
||||
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
|
||||
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
|
||||
thermal throttle T1 3×.**
|
||||
|
||||
### Pool 6 pg_num/pgp_num Mismatch (Fixed 2026-07-24)
|
||||
Pool hdd_disk had pg_num=120, pgp_num=112 (autoscaler reducing to 32).
|
||||
Equalized pg_num to 112. Set nopgchange=true to prevent further autoscaler interference.
|
||||
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
|
||||
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
|
||||
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
|
||||
|
||||
### BlueFS Spillover on osd.8
|
||||
osd.8 spilled 128KiB metadata from db device (2.1GiB of 30GiB) to slow device.
|
||||
Cosmetic warning, no data risk. Fix: `ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-8`
|
||||
|
||||
### Slow Operations on osd.2 and osd.7
|
||||
osd.2 (81% full, fragmentation 0.80) and osd.7 (small HDD) experience slow BlueStore ops.
|
||||
osd.2 NVMe has 92% wear — candidate for replacement.
|
||||
|
||||
### osd.0 NVMe Wear
|
||||
92% Wear, Critical Warning → Austausch planen.
|
||||
|
||||
### EC Pool k=4+m=1 — No Rebalance Headroom
|
||||
With 5 OSDs kein Rebalance Headroom. Siehe Solution Doc: `docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`
|
||||
|
||||
## RBD Management
|
||||
- Proxmox RBD Double-Mount Deadlock Pitfall: Niemals `pct mount` und `pct exec` gleichzeitig auf demselben Container
|
||||
- Siehe Solution Doc: `docs/solutions/bug-fixes/2026-07-23-proxmox-rbd-double-mount-deadlock.md`
|
||||
### worker-04 (VM 139) NotReady in K8s
|
||||
Node offline — not a Ceph issue but affects Ceph CSI attachments.
|
||||
|
||||
## Access
|
||||
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.92`
|
||||
- SSH to Proxmox hosts: `ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.50`
|
||||
- Ceph commands: `ceph status`, `ceph osd tree`, `ceph pg dump pgs`
|
||||
- Mon nodes: proxmox5, proxmox7, proxmox4
|
||||
- Mgr: proxmox5 (active)
|
||||
- Mon nodes: proxmox5 (leader), proxmox4, ms-a2-1, n5pro
|
||||
- Mgr: n5pro (active)
|
||||
|
||||
## Related Skills
|
||||
- `ceph-cluster-administration` (devops)
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
title: "Frigate NVR"
|
||||
category: systems
|
||||
tags: [frigate, nvr, camera, ai, mqtt, proxmox]
|
||||
created: "2026-09-27"
|
||||
modified: "2026-09-27"
|
||||
---
|
||||
|
||||
# Frigate NVR
|
||||
|
||||
> Frigate v0.18.0 auf Proxmox LXC CT151. KI-gestützte Objekterkennung für Kameras Einfahrt + Terrasse.
|
||||
|
||||
## Infrastruktur
|
||||
- **Container:** CT151 auf proxmox6
|
||||
- **IP:** `10.0.30.104:5000`
|
||||
- **Version:** v0.18.0-
|
||||
- **Detector:** OpenVINO CPU (3.2 FPS)
|
||||
- **go2rtc:** v1.9.14
|
||||
- **GPU:** Intel iGPU `/dev/dri/renderD128` (gid=993)
|
||||
|
||||
## Kameras
|
||||
| Name | IP | Substream | Detect FPS |
|
||||
|------|----|-----------|------------|
|
||||
| Einfahrt | `10.0.50.102` | Unterstream | 5.1 fps, 1.0 det |
|
||||
| Terrasse | `10.0.50.103` | Unterstream | 5.0 fps, 2.2 det |
|
||||
|
||||
## MQTT
|
||||
- **Broker:** `10.0.30.10:1883` (Mosquitto auf HA)
|
||||
- **Prefix:** `frigate`
|
||||
- **User:** `frigate`
|
||||
- **Topic:** `frigate/events`
|
||||
|
||||
## Frigate Plus Model
|
||||
- `plus://709bb8ad097786a7f2a37e27684c79a9`
|
||||
- Benötigt `PLUS_API_KEY` (Env-Var in `/config/frigate.env`)
|
||||
|
||||
## Features
|
||||
- Face Recognition (groß): Sarah, Dominik, DHL1, Amazon1, Cleo, Lisa, Annemarie, Marcel Schnitzer, Uli, Eva, Postbote, Fr. Schnitzer
|
||||
- License Plate Recognition (CPU, threshold 0.5)
|
||||
- Semantic Search (groß)
|
||||
- Record: alerts+detections, retain 30 days
|
||||
|
||||
## HA Automation
|
||||
- **ID:** `1714065780832` in `/config/automations.yaml`
|
||||
- **Mode:** `single`, cooldown 600s
|
||||
- **Dedup:** `input_text.frigate_last_event_id` speichert letzten Event
|
||||
- **Flow:** MQTT trigger → 10s delay → snapshot → notify → LLM Vision (optional)
|
||||
- **Labels:** person, car, license_plate, face, dog, cat, amazon, ups, package
|
||||
- **Sublabel-Extraktion:** `sub[0]` (Jinja, Array-Leck-Fix)
|
||||
|
||||
## Bekannte Issues
|
||||
- Stationäre Autos triggerten wiederholt "new" Events → Flood → gefixt mit cooldown+dedup
|
||||
- MQTT ACL blockierte HA-Subscription (User `opendtu` hatte keine Subscribe-Rechte) → gefixt
|
||||
- v0.18 Breaking Changes: `clean_copy` entfernt, `license_plate.mask` Format geändert (list→dict)
|
||||
|
||||
## Old Container
|
||||
- CT120 (frigate): gestoppt, wartet auf Deletion
|
||||
|
||||
## Related
|
||||
- [[systems/homeassistant]] — MQTT, Automations, Notifications
|
||||
- [[systems/noris-ai]] — LLM Vision für Event-Klassifizierung
|
||||
- [[entities/infrastructure]] — CT151 auf proxmox6
|
||||
@@ -35,6 +35,11 @@ modified: "2026-07-24"
|
||||
- VM300/301/302 (Galera): Fluent Bit aktiv
|
||||
- VM310/311 (MaxScale): Fluent Bit aktiv
|
||||
|
||||
## Bekannte Issues
|
||||
- **2026-09-29:** VM302 Zombie-State nach Migration (QEMU running, Services tot, QGA down). Behoben via Stop+Start; SST-Rejoin ~2-3min. Siehe `docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md`
|
||||
- **QEMU-Cmdline-Falle:** `-incoming unix:/run/qemu-server/NNN.migrate -S` bleibt nach Migration im Prozess-String — KEIN Beweis für "paused". Immer `qm monitor <vmid> <<< 'info status'` prüfen.
|
||||
- **HA-Guard:** `ha-manager disable` existiert nicht. Bei Resource-State `request_stop` geht `qm stop` trotzdem durch.
|
||||
|
||||
## Related Skills
|
||||
- `mariadb-galera-cluster-administration` (devops)
|
||||
|
||||
|
||||
+22
-6
@@ -3,18 +3,18 @@ title: Gitea (Git Server + CI)
|
||||
category: systems
|
||||
tags: [gitea, git, ci, actions]
|
||||
created: "2026-07-24"
|
||||
modified: "2026-07-26"
|
||||
modified: "2026-09-27"
|
||||
---
|
||||
|
||||
# Gitea (Git Server + CI)
|
||||
|
||||
## Instanz
|
||||
- **URL:** git.schoen.codes (K8s Ingress, Traefik → gitea-http ClusterIP)
|
||||
- **Internal:** gitea-internal.gitea:3000 (K8s ClusterIP 10.43.5.203)
|
||||
- **SSH:** 10.0.30.202:22 (Cilium LoadBalancer, gitea-ssh service)
|
||||
- **Internal:** gitea-http ClusterIP (headless, `None`) — Port 3000/TCP, via Ingress/Traefik erreichbar
|
||||
- **SSH:** 10.0.30.200:22 (Cilium LoadBalancer, gitea-ssh service, NodePort 31441) — NICHT .202 (alter Wiki-Fehler)
|
||||
- **Version:** 1.27.0 (K8s Helm chart v12.7.0)
|
||||
- **Privat:** NIEMALS öffentlich machen (enthält Secrets)
|
||||
- **CT108 (alte Gitea, 10.0.30.105): DECOMMISSIONED** — gestoppt am 2026-07-24
|
||||
- **CT108 (alte Gitea, 10.0.30.105, v1.25.5): GESTOPPT am 2026-09-17 (genehmigt).** War bis zum Abend NICHT gestoppt trotz älterer Wiki-Claims — Live-Checks (v1.25.5-API-Antwort) schlugen die Papierlage. Vor dem Stop: Full-Census (31 Repos auth, 23 gemeinsam — 22 Tips identisch, dominik/memory K8s-Superset), 29 Bare-Bundle-Archive (143 MB) unter `/home/debian/git-archive/ct108-final/`. Letzter Live-Konsument war die Route `git.familie-schoen.com` im CT99999-Traefik (Upstream .105:3000) — umgebogen auf .203 (K8s), Alias-Host im K8s-Ingress ergänzt (Commit 9e9b6ee). Seither servieren BEIDE Hostnamen v1.27.0.
|
||||
|
||||
## Repositories
|
||||
| Repo | Zweck | Clone |
|
||||
@@ -28,12 +28,15 @@ modified: "2026-07-26"
|
||||
- Ed25519 SSH-Key als Gitea Deploy Key (read-only)
|
||||
- ArgoCD Secret `argocd-repo-k8s-gitea` mit SSH Private Key
|
||||
- `known_hosts` ConfigMap mit Gitea SSH Hostkey
|
||||
- 15 ArgoCD Apps via SSH (`ssh://gitea@git.schoen.codes:22/dominik/iac-homelab.git`)
|
||||
- 17 ArgoCD Apps via SSH (`ssh://git@10.0.30.200:22/dominik/iac-homelab.git`)
|
||||
- 25 ArgoCD Apps gesamt (Stand Sep 2026)
|
||||
|
||||
## Gitea Actions CI
|
||||
- **Primary Runner:** vm200-host (ID 73, act_runner v0.2.13, host backend) — WORKING
|
||||
- **K8s Runner:** gitea-runner deployment in gitea namespace (DinD sidecar + act_runner v0.2.13) — registered but NO TASK ASSIGNMENT (Gitea 1.27.0 bug)
|
||||
- **K8s Runner (gitea NS):** gitea-runner deployment (DinD sidecar + act_runner v0.2.13) — RUNNING, declare erfolgreich (Sep 2026). Vorher 4 Tage Init-crash wegen CSI RBAC Issue (siehe unten).
|
||||
- **K8s Runner (schoenkitchen NS):** gitea-runner StatefulSet (gleiche Architektur) — RUNNING, declare erfolgreich (Sep 2026).
|
||||
- **KNOWN BUG (Gitea 1.27.0):** `actions_ready_job` queue doesn't create `action_task` records for K8s-based runners via gRPC Declare. Only runner 73 (external, pre-registered) receives tasks. Fix: upgrade Gitea to 1.27.1+ or 1.28.x.
|
||||
- **WARNUNG (Sep 2026):** `DEFAULT_ACTIONS_URL = https://gitea.com` ist deprecated in Gitea 1.27.0 — fallback zu `github`. Helm Values setzen nur `ENABLED: true`; der Wert kommt vom Chart-Default. Zu fixen via Helm values override: `DEFAULT_ACTIONS_URL: self`.
|
||||
- **No runner management API** in Gitea 1.27.0 (`/api/v1/admin/runners` → 404). Manage via DB: `gitea.action_runner` table, `is_disabled` column, `deleted` for soft-delete.
|
||||
- Workflow: `.gitea/workflows/rebuild-infrastructure.yml`
|
||||
- Achtung: `actions/checkout@v4` versucht github.com zu erreichen — muss gemapped werden
|
||||
@@ -51,6 +54,19 @@ modified: "2026-07-26"
|
||||
- **Verwendung:** `GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o StrictHostKeyChecking=no" git push ssh://git@10.0.30.202:22/dominik/iac-homelab.git HEAD:main`
|
||||
- **Wichtig:** Gitea HTTP (Port 3000) hat keine externe IngressRoute — SSH (Port 22 via LoadBalancer) ist der einzige zuverlässige Weg für Git-Pushes von außerhalb K8s.
|
||||
|
||||
## Known ArgoCD Issues (Sep 2026)
|
||||
- **`gitea` app:** Sync=Succeeded, Health=Progressing — Ingress hat keine LoadBalancer IP (Traefik reporting gap). Kosmetisch, Funktionalität OK.
|
||||
- **`gitea-config` app:** OutOfSync — `Deployment gitea-runner` driftet (manuelle Änderung am Image?). Sync würde helfen.
|
||||
- **`immich` + `immich-config` apps:** OutOfSync — zwei Apps zeigen auf überlappende Pfade (`rendered` vs roh) im iac-homelab Repo. Architektur-Issue, kein Gitea-Problem.
|
||||
- **`authelia` app:** Health=Progressing — gleiche Ingress-LB-IP Thematik wie gitea.
|
||||
|
||||
## Incident 2026-09-27: CSI RBAC → Gitea Pod Init-Crash
|
||||
- **Symptom:** Gitea Pod 4+ Tage in `Init:0/3`, schoenkitchen runner in `PodInitializing`.
|
||||
- **Root Cause:** Ceph CSI RBD provisioner `csi-attacher` hatte keine Berechtigung, `csinodes` zu lesen → VolumeAttachment blieb pending → PVC konnte nicht mounten.
|
||||
- **Fix:** RBAC ClusterRole `ceph-rbd-external-attacher-runner` um `storage.k8s.io/csinodes` Resource ergänzt, provisioner Pods neu gestartet.
|
||||
- Nach Fix: Alle 24 VolumeAttachments `true`, Gitea Pod Running nach Pod-Delete.
|
||||
- **Lesson:** Bei CSI-basierten PVs immer zuerst VolumeAttachment Status prüfen, nicht nur PVC Phase.
|
||||
|
||||
## Related
|
||||
- [[systems/rke2-kubernetes]]
|
||||
- [[concepts/gitops-workflow]]
|
||||
|
||||
@@ -10,8 +10,10 @@ modified: "2026-07-24"
|
||||
|
||||
## Deployment
|
||||
- **Namespace:** hindsight (K8s)
|
||||
- **API:** LoadBalancer `10.0.30.208:9177` (**NICHT localhost**)
|
||||
- **API:** LoadBalancer `10.0.30.208:9177` (**NICHT localhost**, NICHT .201 — dort sitzt seit Rebuild 2026-08-01 Grafana!)
|
||||
- **Health:** `curl -s http://10.0.30.208:9177/health`
|
||||
- **Port-Mapping:** Svc 9177 → Container 8888; NodePort-Fallback 31577 (jeder Node)
|
||||
- **LB-IP-Drift-Warnung:** LB-IPs verschoben sich beim Rebuild 2026-08-01 (hindsight .201→.208). Bei Health-Failure IMMER zuerst `kubectl -n hindsight get svc hindsight-api` gegenprüfen, statt Referenz-IP zu vertrauen
|
||||
- **Backend:** PostgreSQL + pgvector
|
||||
- **Config:** `~/.hermes/hindsight/config.json` (api_url gesetzt)
|
||||
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
---
|
||||
title: "Home Assistant"
|
||||
category: systems
|
||||
tags: [homeassistant, smart-home, mqtt, automation]
|
||||
created: "2026-09-27"
|
||||
modified: "2026-09-27"
|
||||
---
|
||||
|
||||
# Home Assistant
|
||||
|
||||
> Smart Home Zentrale auf Proxmox. Steuerung, Automatisierung und Benachrichtigungen.
|
||||
|
||||
## Infrastruktur
|
||||
- **Host:** `10.0.30.10` (HAOS VM)
|
||||
- **SSH:** `hassio@10.0.30.10` (PW: 1P Vault "Hermes")
|
||||
- **URL:** `https://homeassistant.familie-schoen.com`
|
||||
- **Container:** `homeassistant` (Docker)
|
||||
|
||||
## MQTT
|
||||
- **Broker:** Mosquitto Add-on v7.1.1 auf localhost:1883
|
||||
- **HA MQTT User:** ehemals `opendtu` (BROKEN — ACL blockiert), gefixt
|
||||
- **ACL:** `/etc/mosquitto/acl` definiert `user homeassistant` + `user addons`
|
||||
- **Auth Plugin:** `go-auth.so` (files,http backends)
|
||||
|
||||
## Automations
|
||||
- **File:** `/config/automations.yaml` (35 Automations)
|
||||
- **Schreibmethode:** SSH → `docker exec homeassistant chmod 666`, danach restore 644
|
||||
- **Reload:** REST API mit JWT (HS256, signed from `/config/.storage/auth`)
|
||||
|
||||
### Frigate Einfahrt Notification (ID 1714065780832)
|
||||
- MQTT trigger `frigate/events` → filter `type=='new'` + cameras [Einfahrt,Terrasse] + labels [person,car,...]
|
||||
- Mode: `single`, cooldown 600s
|
||||
- Dedup: `input_text.frigate_last_event_id`
|
||||
- Flow: delay 10s → snapshot → notify iPhone → optional LLM Vision
|
||||
- Siehe [[systems/frigate]]
|
||||
|
||||
## Notify Services
|
||||
| Service | Status |
|
||||
|---------|--------|
|
||||
| `notify.mobile_app_iphone_dominik` | ✅ aktiv |
|
||||
| `notify.mobile_app_sarahs_iphone_app` | verfügbar |
|
||||
| `notify.mobile_app_ipad_2` | verfügbar |
|
||||
| `notify.mobile_app_sm_x205` | verfügbar |
|
||||
|
||||
## Entitäten
|
||||
- Kameras: `camera.einfahrt_2`, `camera.terrasse_2` (Suffix `_2` wegen verwaister Integration)
|
||||
- Motion: `binary_sensor.einfahrt_motion_2`, `binary_sensor.terrasse_motion_2`
|
||||
- Input Text: `input_text.frigate_last_event_id` (dedup storage, max 255 chars)
|
||||
|
||||
## Snapshots
|
||||
- Gespeichert: `/config/www/snapshots/{camera}_latest.jpg`
|
||||
- URL: `/local/snapshots/` (HTTP 200, keine Auth)
|
||||
|
||||
## JWT Auth
|
||||
1. `jwt_key` aus `/config/.storage/auth` lesen
|
||||
2. Client `Hermes_202606` (token id `1444b2c6757b4d66a6f8f8e5899b4e6a`)
|
||||
3. HS256 signieren, Bearer Header
|
||||
4. `?return_response=true` für Service-Call Responses
|
||||
|
||||
## Integrations
|
||||
- Frigate (MQTT)
|
||||
- LLM Vision (noris AI `gemma-4-31b-it`)
|
||||
- Tibber (Strom)
|
||||
- Marstek Speicher (VENUS-E, IP 10.0.50.113)
|
||||
- Various sensors (Xiaomi BLE, etc.)
|
||||
|
||||
## Related
|
||||
- [[systems/frigate]] — NVR Integration
|
||||
- [[systems/noris-ai]] — LLM Vision Provider
|
||||
- [[reference/ip-map]] — IP Assignments
|
||||
@@ -0,0 +1,85 @@
|
||||
# Laya Decision Model
|
||||
|
||||
> 421M param ModernBERT-large Classifier (Apache 2.0) auf CT152, CPU-only.
|
||||
> Nutze in normalen Sessions fuer schnelle Klassifizierung, Binaerentscheidungen und Pre-Filter.
|
||||
|
||||
## Zugang
|
||||
- **Endpoint:** `POST http://10.0.30.152:8000/predict`
|
||||
- **Payload:** `{"state": "<text>", "questions": {...}}`
|
||||
- **Health:** `GET http://10.0.30.152:8000/health`
|
||||
- **Presets:** `GET http://10.0.30.152:8000/presets/{router|guard|moderation|triage}`
|
||||
|
||||
## Entscheidungstypen (Primitives)
|
||||
|
||||
| Typ |用途 | Return Fields |
|
||||
|-----|------|---------------|
|
||||
| `choice` | Klassifizierung in N Labels | `choice`, `answer_confidence`, `probabilities` |
|
||||
| `noul` | Ja/Nein mit Wahrscheinlichkeit | `noul` (0..1), `answer_confidence` |
|
||||
| `score` | Ordinale Bewertung | `score`, `answer_confidence` |
|
||||
|
||||
## Verwendung in Sessions
|
||||
|
||||
**Praeferieren fuer:**
|
||||
- Pre-Filter vor teuren LLM-Calls (z.B. "ist diese Email eine Rechnung?" → nur bei "ja" LLM aufrufen)
|
||||
- Binaerentscheidungen: alert/skip, escalate/ignore, move/keep
|
||||
- Multi-Kategorie-Klassifizierung mit Confidence
|
||||
- Gatekeeping: Notification-Suppression, Alert-Filtering
|
||||
|
||||
**NICHT geeignet fuer:**
|
||||
- Textgenerierung / Zusammenfassungen (dafür LLM verwenden)
|
||||
- Komplexe Reasoning-Tasks
|
||||
- Embeddings / Semantische Suche (dafür Harrier/Hindsight)
|
||||
|
||||
## Example Call
|
||||
|
||||
```python
|
||||
import json, urllib.request
|
||||
|
||||
payload = json.dumps({
|
||||
"state": "Von: amazon.de\nBetreff: Bestellbestätigung #12345",
|
||||
"questions": {
|
||||
"kategorie": {
|
||||
"type": "choice",
|
||||
"instructions": "Welche Kategorie?",
|
||||
"criteria": {
|
||||
"rechnung": "Rechnung, Invoice",
|
||||
"bestellung": "Bestellbestätigung, Order",
|
||||
"werbung": "Newsletter, Marketing"
|
||||
}
|
||||
},
|
||||
"ignorieren": {
|
||||
"type": "noul",
|
||||
"instructions": "Soll ignoriert werden?"
|
||||
}
|
||||
}
|
||||
}).encode()
|
||||
|
||||
req = urllib.request.Request(
|
||||
"http://10.0.30.152:8000/predict",
|
||||
data=payload,
|
||||
headers={"Content-Type": "application/json"}
|
||||
)
|
||||
result = json.loads(urllib.request.urlopen(req, timeout=30).read())
|
||||
# result["answers"]["kategorie"]["choice"] → "bestellung"
|
||||
# result["answers"]["kategorie"]["answer_confidence"] → 0.99
|
||||
```
|
||||
|
||||
## Performance
|
||||
- Latenz: ~1.5-1.7s warm cache (CPU)
|
||||
- Load time: ~18s (cold start)
|
||||
- systemd service: `laya.service` (enabled, onboot)
|
||||
|
||||
## Einsatzgebiete (aktiv)
|
||||
- **Rechnungen-Organizer** (Cron `f773f8c23230`): 12-Kategorie Email-Klassifizierung, 3-Schichten-Safety
|
||||
- Weitere Kandidaten: Beleg-Sammler, SRE Network Recon, Backup Digest, Frigate Event Gate
|
||||
|
||||
## Constraints
|
||||
- 421M Modell → niedrige Confidence bei ambiguous Inputs (Feature, nicht Bug!)
|
||||
- Batch funktioniert nicht — ein Request pro Input-Instanz
|
||||
- `noul` ist nicht zuverlaessig fuer kritische Entscheidungen allein — immer mit `choice` kombinieren
|
||||
- Confidence-Threshold empfohlen (≥0.6 fuer Moves, ≥0.5 fuer Ignores)
|
||||
|
||||
## Related
|
||||
- [[concepts/email-organization]] — Rechnungs-Organizer Architecture
|
||||
- [[entities/infrastructure]] — CT152 auf proxmox6
|
||||
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
title: "noris AI Platform"
|
||||
category: systems
|
||||
tags: [ai, llm, noris, gpu, embeddings]
|
||||
created: "2026-09-27"
|
||||
modified: "2026-09-29"
|
||||
---
|
||||
|
||||
# noris AI Platform (ai.noris.de)
|
||||
|
||||
> Interne AI-Plattform der noris Network AG. Bereitstellung von LLMs, Embeddings und Image Generation.
|
||||
|
||||
## Endpoints
|
||||
- **Chat:** `https://ai.noris.de/v1/chat/completions`
|
||||
- **Embeddings:** `https://ai.noris.de/v1/embeddings`
|
||||
- **Images:** ⚠️ `/v1/images/generations` wird vom Bifrost Gateway **NICHT** unterstützt. Image-Gen-Modelle (qwen-image-2-1) werden über `/v1/chat/completions` angesprochen — das Bild kommt als base64-PNG im `content`-Array zurück (Typ `image_url`, `data:image/png;base64,...`). Siehe `references/vllm-image-generation.md` im Skill `serving-llms-vllm`.
|
||||
|
||||
## Modelle
|
||||
| Typ | Modell-ID | Hinweise |
|
||||
|-----|-----------|----------|
|
||||
| Flagship LLM | `glm-5-2` | Primary, OpenRouter-kompatibel |
|
||||
| General | `gemma-4-31b-it` | Vision-fähig, genutzt von LLM Vision |
|
||||
| Large MoE | `gpt-oss-120b` | |
|
||||
| Mid-range | `qwen3.6-27b` | |
|
||||
| Mid-range | `qwen3.8-27b` | |
|
||||
| Fast | `ds-v4-flash` | Low-latency, Paperless OCR |
|
||||
| Embedding | `harrier` | Vektorembeddings |
|
||||
| Image Gen | `qwen-image-2-1` | Via `/v1/chat/completions` (NOT images/generations). Base64-PNG im content-Array. ~30s/ Bild. |
|
||||
|
||||
## Verbraucher
|
||||
- **Hermes Agent** — Primärmodell `glm-5-2` via OpenRouter
|
||||
- **HA LLM Vision** — `gemma-4-31b-it` für Bildanalyse (Frigate Events)
|
||||
- **Paperless** — `ds-v4-flash` für OCR/Kategorisierung
|
||||
- **Personal Coach Bot** — `glm-5-2` via noris direkt
|
||||
- **Dynamic Coach** — `glm-5-2` via noris direkt, `qwen-image-2-1` für Visualisierungen
|
||||
|
||||
## Related
|
||||
- [[systems/frigate]] — nutzt noris AI für Event-Klassifizierung
|
||||
- [[systems/homeassistant]] — LLM Vision Integration
|
||||
- [[systems/paperless]] — OCR via ds-v4-flash
|
||||
@@ -0,0 +1,34 @@
|
||||
---
|
||||
title: "Paperless-ngx"
|
||||
category: systems
|
||||
tags: [paperless, documents, oidc, ocr]
|
||||
created: "2026-09-27"
|
||||
modified: "2026-09-27"
|
||||
---
|
||||
|
||||
# Paperless-ngx
|
||||
|
||||
> Dokumentenmanagement mit OCR, OIDC-Login und AI-Kategorisierung.
|
||||
|
||||
## Zugriff
|
||||
- **URL:** `https://dokumente.familie-schoen.com`
|
||||
- **mTLS:** `https://dokumente-mtls.familie-schoen.com` (auto-login als `dominik`)
|
||||
- **PKCS12:** `dominik-dokumente-mtls.p12` (PW: siehe 1P Vault "Hermes")
|
||||
|
||||
## Auth
|
||||
- OIDC via Authelia (`dominik@schoen.eu`)
|
||||
- Break-Glass lokaler User: `dominik`
|
||||
- mTLS Client Cert: CN=dominik, gültig bis Juli 2028
|
||||
|
||||
## Konfiguration
|
||||
- **AI Backend:** `ds-v4-flash@ai.noris.de` (noris AI)
|
||||
- **Memory:** 4Gi (PAT-002)
|
||||
- **Mail Import:** `dokumente@familie-schoen.com` (iCloud mailbox, max 30 Tage)
|
||||
- **Owner:** dominik
|
||||
|
||||
## Known Issue
|
||||
- PAT-002: Memory-Limit 4Gi erforderlich, sonst OOM bei großen OCR-Batches
|
||||
|
||||
## Related
|
||||
- [[concepts/credential-policy]] — 1Password, mTLS Zertifikate
|
||||
- [[systems/noris-ai]] — ds-v4-flash für OCR
|
||||
@@ -3,14 +3,14 @@ title: Proxmox VE Cluster
|
||||
category: systems
|
||||
tags: [proxmox, virtualization, lxc, qemu, pve]
|
||||
created: "2026-04-28"
|
||||
modified: "2026-07-24"
|
||||
modified: "2026-09-30"
|
||||
---
|
||||
|
||||
# Proxmox VE Cluster
|
||||
|
||||
## Cluster-Konfiguration
|
||||
- **Version:** PVE 9.2.10, Kernel 7.0.14-12-pve (upgraded 2026-08-17)
|
||||
- **Nodes:** 9 (Quorum OK)
|
||||
- **Version:** PVE 9.2.20, Kernel 7.0.14-19-pve (upgraded 2026-09-25)
|
||||
- **Nodes:** 8 (Quorum OK, proxmox2 dauerhaft entfernt)
|
||||
- **Hypervisoren:** 10.0.20.x
|
||||
- **Guests:** ~30 LXC + ~10 QEMU VMs
|
||||
|
||||
@@ -50,13 +50,92 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
||||
- Benötigte modprobe.d Config:
|
||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||
- Worker-05 (VM 102) läuft auf ms-a2-2 mit funktionierendem GPU-Passthrough
|
||||
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
|
||||
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
|
||||
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
|
||||
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
|
||||
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
|
||||
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
|
||||
|
||||
## Bekannte Probleme
|
||||
- osd.0 NVMe 92% Wear — Austausch planen
|
||||
- osd.2/5 nearfull (93-94%) — entlasten
|
||||
- CT110 kaputte libc — Reparatur ausstehend
|
||||
- ms-a2-1 GPU-Passthrough: Config gefixt, Reboot zur Verifikation ausstehend
|
||||
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
||||
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
||||
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
|
||||
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
|
||||
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
|
||||
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
|
||||
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
|
||||
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
|
||||
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
|
||||
|
||||
## HA Rules (PVE 9.2 Rules System)
|
||||
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||
Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
||||
|
||||
**Resource-Affinity (negative = anti-collocation):**
|
||||
|
||||
*RKE2 Control Plane (etcd-Quorum braucht 2/3):*
|
||||
- `rke2-cp-anti-112-122`: vm:112 ↔ vm:122 — nie auf gleichem Host
|
||||
- `rke2-cp-anti-112-126`: vm:112 ↔ vm:126 — nie auf gleichem Host
|
||||
- `rke2-cp-anti-122-126`: vm:122 ↔ vm:126 — nie auf gleichem Host
|
||||
|
||||
*RKE2 Worker:*
|
||||
- `rke2-worker-anti-102-128`: vm:102 ↔ vm:128 — nie auf gleichem Host
|
||||
- `rke2-worker-anti-102-139`: vm:102 ↔ vm:139 — nie auf gleichem Host
|
||||
- `rke2-worker-anti-128-139`: vm:128 ↔ vm:139 — nie auf gleichem Host
|
||||
|
||||
*Galera:*
|
||||
- `galera-anti-300-301`: vm:300 ↔ vm:301 — nie auf gleichem Host
|
||||
- `galera-anti-300-302`: vm:300 ↔ vm:302 — nie auf gleichem Host
|
||||
- `galera-anti-301-302`: vm:301 ↔ vm:302 — nie auf gleichem Host
|
||||
|
||||
*MaxScale:*
|
||||
- `maxscale-anti-310-311`: vm:310 ↔ vm:311 — nie auf gleichem Host
|
||||
|
||||
**Node-Affinity (non-strict, failover allowed):**
|
||||
|
||||
*RKE2 CP:*
|
||||
- `na-vm112`: vm:112 → proxmox3, proxmox5, proxmox7
|
||||
- `na-vm122`: vm:122 → ms-a2-1, ms-a2-2
|
||||
- `na-vm126`: vm:126 → proxmox4, proxmox5, proxmox6
|
||||
|
||||
*RKE2 Worker:*
|
||||
- `na-vm102`: vm:102 → proxmox5, proxmox6, proxmox7 (NICHT ms-a2-2!)
|
||||
- `na-vm128`: vm:128 → ms-a2-2, proxmox5, proxmox7
|
||||
- `na-vm139`: vm:139 → n5pro, proxmox3, proxmox4
|
||||
|
||||
*Hermes:*
|
||||
- `na-vm230`: vm:230 → n5pro, proxmox5, proxmox6
|
||||
|
||||
*Galera/MaxScale:*
|
||||
- `na-vm300`: vm:300 → n5pro, proxmox3, proxmox4
|
||||
- `na-vm301`: vm:301 → proxmox6, proxmox5, proxmox7
|
||||
- `na-vm302`: vm:302 → ms-a2-2, ms-a2-1
|
||||
- `na-vm310`: vm:310 → proxmox7, proxmox5, proxmox4
|
||||
- `na-vm311`: vm:311 → ms-a2-1, ms-a2-2
|
||||
|
||||
**Aktuelle Verteilung (alle Anti-Collocation erfüllt):**
|
||||
| Role | VM | Node |
|
||||
|------|----|------|
|
||||
| RKE2 CP-01 | 112 | proxmox3 |
|
||||
| RKE2 CP-02 | 122 | ms-a2-1 |
|
||||
| RKE2 CP-03 | 126 | proxmox4 |
|
||||
| RKE2 Worker-01 | 128 | proxmox5 |
|
||||
| RKE2 Worker-04 | 139 | n5pro |
|
||||
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
||||
| Hermes-Agent-01 | 230 | n5pro |
|
||||
| Galera db1 | 300 | n5pro |
|
||||
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
|
||||
| Galera db3 | 302 | ms-a2-2 |
|
||||
| MaxScale-01 | 310 | proxmox7 |
|
||||
| MaxScale-02 | 311 | ms-a2-1 |
|
||||
|
||||
> ⚠️ PVE 9.2 Constraints:
|
||||
> - Resources in resource-affinity rules dürfen keine multi-priority node-affinity haben (gleiche Priorität für alle Nodes erforderlich).
|
||||
> - `ha-manager add` MUSS vor `ha-manager rules add` kommen — sonst "cannot use unmanaged resource".
|
||||
> - Bei gleichzeitigem HA-Add + Anti-Collocation-Violation kann HA-Manager deadlocks (beide VMs auf `migrate` fest). Lösung: eine VM temporär aus HA entfernen, manuell migrieren, dann re-add.
|
||||
> - Online-Migration von VMs mit hohen Memory-Writes (>12GB dirty pages) kann `broken pipe` fehlschlagen. Offline-Migration (stop→migrate→start) als Fallback.
|
||||
|
||||
## Related
|
||||
- [[systems/ceph-cluster]]
|
||||
|
||||
@@ -3,7 +3,7 @@ title: RKE2 Kubernetes Cluster
|
||||
category: systems
|
||||
tags: [kubernetes, rke2, cilium, argocd, cnpg, gitops]
|
||||
created: "2026-07-24"
|
||||
modified: "2026-07-24"
|
||||
modified: "2026-09-17"
|
||||
---
|
||||
|
||||
# RKE2 Kubernetes Cluster
|
||||
@@ -22,16 +22,21 @@ modified: "2026-07-24"
|
||||
| cp-03 | 10.0.30.53 | Control Plane |
|
||||
| worker-01 | 10.0.30.63 | Worker |
|
||||
| worker-04 | 10.0.30.64 | Worker |
|
||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2; 30.09. OOM-Zombie-Freeze behoben, siehe PAT-010) |
|
||||
|
||||
## Storage
|
||||
- **Ceph CSI**: ceph-flash (fast), ceph-hdd (bulk)
|
||||
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`
|
||||
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
||||
- **⚠️ StorageClass IaC (2026-09-29):** `ceph-flash` + `ceph-hdd-replica` waren MANUELL erstellt (nicht in Git). Jetzt committed unter `clusters/main/storage/`. Alle RBD SCs MÜSSEN `controller-expand-secret-name/namespace` haben — fehlt dies, schlagen Volume Expansions fehl ("provided secret is empty") und triggern Retry-Loops die API Server überlasten.
|
||||
- **⚠️ PVs snapshotten SC-Parameter zur Provisionierungszeit** — SC fixen reicht NICHT; bestehende PVs brauchen individuellen Patch mit `controllerExpandSecretRef` falls sie ohne erstellt wurden.
|
||||
- **⚠️ StorageClass `parameters` sind IMMUTABLE** — können nicht gepatched werden, müssen gelöscht+neu erstellt werden.
|
||||
- **⚠️ Snapshot-CRD-Flavor (seit Rebuild 2026-08-01):** `volumesnapshotclasses`-CRD (RKE2-Addon `rke2-snapshot-controller-crd`) hat **FLAT-Schema** — `driver`/`deletionPolicy`/`parameters` auf TOP-LEVEL, `spec:` existiert nicht im Schema (Upstream-CRD wäre nested!). Niemals struktur-intuitiv "reparieren" — CRD-Schema lesen. Live-VSCs: `ceph-rbd-snapclass` (default) + `cephfs-snapclass`
|
||||
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`; Specs **100Gi data / 20Gi WAL** (grow-only — Shrink wird von CNPG-Admission verboten, Git immer nach oben alignieren)
|
||||
- **⚠️ etcd auf Ceph RBD (~20ms WAL fsync auf allen CP-Nodes)** — architektonisches Risiko; unter API Server Write Pressure → gRPC DeadlineExceeded Cascade. Lokales NVMe für etcd data dirs empfohlen.
|
||||
|
||||
## GitOps
|
||||
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
|
||||
- 15 Applications via SSH (`ssh://git@10.0.30.202:22/dominik/iac-homelab.git`)
|
||||
- **Gitea SSH user is `git`, not `gitea`** — `git.schoen.codes` resolves to Traefik (10.0.30.208, HTTP only), SSH is on 10.0.30.202:22
|
||||
- **Feed-Quelle (Stand 2026-09-17):** `http://10.0.30.105:3000/...` = Legacy-CT108-Gitea (noch aktiv!). Geplanter Flip auf `ssh://git@10.0.30.200:22/dominik/iac-homelab.git` (end-to-end verifiziert), danach CT108-Stop (Freigabe Dominik)
|
||||
- **Gitea SSH user is `git`, not `gitea`** — SSH-LoadBalancer = **10.0.30.200:22** (svc gitea-ssh, targetPort 2222, NodePort 31441). Legacy-.202 ist TOD (Timeout). Repo-Historien am 2026-09-17 konsolidiert (identische Tips auf beiden Remotes, Merge `d76d3e8` + `8c741c5`)
|
||||
- Workflow: IaC repo auschecken → ändern → commit+push → ArgoCD sync → verify → lokale Kopie löschen
|
||||
- Siehe [[concepts/gitops-workflow]]
|
||||
|
||||
@@ -61,6 +66,19 @@ modified: "2026-07-24"
|
||||
- IaC: unified `amd-gpu` PCI mapping (n5pro + ms-a2-1 + ms-a2-2) + dynamic hostpci (one worker block, gpu flag)
|
||||
- ms-a2-1 has kernel BUG with 1002:13c0 (renderD128 missing) — worker-05 moved to ms-a2-2 (identical hardware)
|
||||
|
||||
## Capacity Management
|
||||
- **Descheduler** v0.30.0 deployed als ArgoCD-App (`clusters/main/descheduler/manifests.yaml`)
|
||||
- Plugin: `LowNodeUtilization` (thresholds 30/50/20, targetThresholds 70/70/60)
|
||||
- Interval: 2m (`--descheduling-interval=2m`)
|
||||
- `nodeFit: true` auf Profile-Ebene (nicht Plugin-Ebene — v0.30 Schema!)
|
||||
- RBAC: benötigt `list/watch` auf `namespaces` (fehlt in Default-Helm-Chart)
|
||||
- Raw Manifests statt Helm Chart (Chart 0.30 hat params/args-Bug, Chart 0.35 hat falsche API-Version)
|
||||
- Siehe Solution Doc `2026-09-30-k8s-cp01-relief-descheduler.md`
|
||||
- **CP-01 Relief (2026-09-30):** hindsight-api + hindsight-postgres + paperless von cp-01 → worker-05 migriert
|
||||
- cp-01 RAM: 93% → 43%; worker-05 bei ~28%
|
||||
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren
|
||||
- Live-only nodeSelector (hindsight-api, Sep-17-Hotfix) war nie in Git → Live-Patch nötig
|
||||
|
||||
## Known Pitfalls
|
||||
- `enableServiceLinks: false` bei Apps deren Service-Name mit Env-Vars kollidiert (z.B. Paperless `PAPERLESS_PORT`)
|
||||
- ArgoCD `--force` kann nicht mit ServerSideApply kombiniert werden
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
---
|
||||
title: Sarah-Hermes (VM107)
|
||||
category: systems
|
||||
tags: [hermes, webui, vm, sarah, telegram, backup]
|
||||
created: "2026-09-30"
|
||||
modified: "2026-09-30"
|
||||
related: [systems/rke2-kubernetes, reference/ip-map, systems/noris-ai]
|
||||
---
|
||||
|
||||
# Sarah-Hermes (VM107)
|
||||
|
||||
Zweite, vollständig isolierte Hermes-Instanz für Sarah (Allround-Assistentin:
|
||||
Erinnerungen, Planung, Smalltalk). Getrennte Memories/Sessions/Skills/Keys von
|
||||
den 5 Owner-Profilen.
|
||||
|
||||
## Eckdaten
|
||||
|
||||
| Attribut | Wert |
|
||||
|----------|------|
|
||||
| VM-ID | 107 (auto-allokiert) |
|
||||
| Node | n5pro (Template 9000 debian-12-cloudinit) |
|
||||
| IP | 10.0.30.66/24 (static via DHCP reservation) |
|
||||
| Specs | 2 vCPU / 4 GB RAM / 32 GB Disk (vm_disks/RBD) |
|
||||
| Access | SSH `debian@10.0.30.66` mit `~/.ssh/id_ed25519_cloudinit` |
|
||||
| Tofu | `iac-homelab/epic-8-sarah-hermes/tofu/` |
|
||||
| Ansible | `iac-homelab/epic-8-sarah-hermes/ansible/` |
|
||||
| Plan | `iac-homelab/docs/plans/2026-09-30-sarah-hermes-vm.md` |
|
||||
|
||||
## Stack
|
||||
|
||||
- **hermes-webui** (ghcr.io/nesquena/hermes-webui:latest), Single-Container,
|
||||
Port 8787 (0.0.0.0 gebunden, UFW erlaubt nur LAN), Password-Auth.
|
||||
Compose: `/home/debian/hermes-webui/docker-compose.yml` auf der VM.
|
||||
- **Agent-Runtime:** nousresearch/hermes-agent geklont nach
|
||||
`~/.hermes/hermes-agent` auf der VM; installiert in `/app/venv` im Container
|
||||
(editable). **Wichtig:** Core-Deps im pyproject sind hinter
|
||||
`python_version >= '3.14'`-Markern gepinnt → auf Python 3.12 installiert
|
||||
`-e .` NULL Deps. Manual-Dep-Bootstrapping nötig (siehe Pitfalls).
|
||||
- **LLM:** noris-Provider (`https://ai.noris.de/v1`, Default-Modell
|
||||
`vllm/release/glm-5-2`), Key via `HERMES_CUSTOM_NORIS_API_KEY` aus
|
||||
`.env` (Compose mapped explizit in den Container).
|
||||
- **Telegram:** geplant (blockiert auf BotFather-Token von Sarah/Dominik).
|
||||
Bei Aktivierung: `gateway.telegram_enabled: true` + Token in `.env`;
|
||||
Webhook/API-Server-Ports bleiben disabled (Konfliktvermeidung).
|
||||
|
||||
## Secrets (1Password, Vault: Hermes)
|
||||
|
||||
- `sarah-hermes-webui` → HERMES_WEBUI_PASSWORD
|
||||
- `sarah-hermes-noris-key` → HERMES_CUSTOM_NORIS_API_KEY
|
||||
|
||||
## Backup
|
||||
|
||||
- Daily vzdump-Job `backup-2e8a34e3-66cb` (23:00, `all=1`, exclude 301,302,310,311,137,147,501)
|
||||
→ **deckt VM107 ab** (Storage `noris_v4` = PBS Datastore `noris` @ 10.0.30.119).
|
||||
- Manueller Verify-Lauf am 30.09.: TASK OK in 48s (inkrementell, 88% reuse).
|
||||
- Offsite: folgt dem regulären `push-offsite` Sync-Job (Pull↔Push-Korrektur
|
||||
vom 19.09.).
|
||||
|
||||
## Known Issues / Pitfalls
|
||||
|
||||
1. **Python-Version-Mismatch:** hermes-agent pyproject pins Core-Deps an
|
||||
`python_version >= '3.14'`; WebUI-Container läuft auf 3.12 → `-e .`
|
||||
installiert keine Deps. Fix: manuell `pip install` der gepinschten Pakete
|
||||
+ iterativer Missing-Import-Loop bis `import run_agent` klappt.
|
||||
2. **hermes update Ownership-Konflikt:** `hermes update`-Completion beschwert
|
||||
sich über uid 0 vs. uid 1000 auf `/app/venv/bin/hermes-acp`. Kosmetisch,
|
||||
Betriebsbetrieb unbeeinträchtigt. Fix-Idee: `chown` im Entry-Point.
|
||||
3. **telegram-bridge Notify-Target defekt** (seit 19.09., Connection refused):
|
||||
betrifft vzdump-Notifications clusterweit, nicht nur VM107. Separater Fix.
|
||||
|
||||
## Verification History
|
||||
|
||||
- 2026-09-30: Deploy + E2E-Test (Login 200, Agent-Antwort "HALLO" via glm-5-2).
|
||||
- 2026-09-30: Manueller vzdump → TASK OK (48s, inkrementell).
|
||||
Reference in New Issue
Block a user