Files
memory/log.md
T

292 lines
29 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Memory Log
## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur
- Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen.
- Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate.
- SMART-Korrektur: Kingston SFYRDK2000G (osd.2-Träger) ist GESUND (wear 6 %, media_errors 0, spare 100 %, PoH 1469) — frühere "Austausch-Kandidatin"-These RETRAKIERT. Vorfall war Controller-/Fabric-Ebene. Neu beobachten: Temps 70/78 °C, T1-throttle 3×.
- PAT-014 erweitert (Smart-Log-Abschnitt), ceph-cluster.md korrigiert, Commit cc2f237.
## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014)
- osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin. *(Korrektur später am selben Tag: SMART-Audit → gesund, kein Austausch; siehe nächster Eintrag.)*
- osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend.
- ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben.
- Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014.
## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013)
- Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün.
- Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok.
- **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate.
- Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit <SHA>.
## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012)
- Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok.
- Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom).
- Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172.
## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011)
- Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON).
- Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200.
- Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body.
- Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales).
## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen
- proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%.
- Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets.
- VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP.
- Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt.
## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet
- GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md.
- Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc.
- proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen.
- Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert.
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX
- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared
- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr
- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return
- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010)
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
- E2E verifiziert: Login 200, Agent-Antwort via vllm/release/glm-5-2 @ ai.noris.de
- PITFALL: hermes-agent pyproject pinnt Core-Deps hinter python_version>='3.14'-Markers → auf Python 3.12 installiert `pip install -e .` NULL Deps ("AIAgent not available"). Fix: manuelle Pin-Installation + Missing-Import-Loop bis `import run_agent`
- Backup: täglich 23:00 Job backup-2e8a34e3-66cb (all=1) deckt VM107; manueller Verify TASK OK 48s
- Wiki: systems/sarah-hermes.md, ip-map.md erweitert; iac-homelab commits 0d5f017/1ce4816
- OFFEN: Telegram-Bot blockiert auf BotFather-Token (Sarah/Dominik)
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
- Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min)
- Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05
- Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2)
- Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes)
- Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
- Gitea Pod + schoenkitchen runner durch Pod-Delete wiederhergestellt
- ArgoCD Audit: 17/25 Apps via Gitea SSH, 20/25 Synced+Healthy
- Known Issues: `gitea`/`authelia` Health=Progressing (Ingress-LB-IP Gap), `immich` doppelt verwaltet, `gitea-config` Deployment Drift
- `DEFAULT_ACTIONS_URL=https://gitea.com` deprecated → Helm override auf `self` empfohlen
- Wiki `systems/gitea.md` aktualisiert: Runner Status, Incident, ArgoCD Issues
- Solution Doc: `docs/solutions/bug-fixes/2026-09-27-gitea-csi-rbac-volumeattachment-fix.md`
## [2026-09-27] architecture | Memory Layer Restructuring + Laya Session-Nutzung
- USER.md bereinigt: Infra-Fakten entfernt, nur noch User-Preferences/Safety-Rules (1.087/1.375 chars)
- MEMORY.md ausgedünnt: 2.145→1.664 chars, alle Infra-Details zeigen auf Wiki-Seiten mit `→Wiki` Pointern
- 4 neue Wiki-Seiten: systems/noris-ai, systems/frigate, systems/homeassistant, systems/paperless
- Neues Concept: concepts/memory-layer-architecture — Entscheidungsbaum "was wohin gehört"
- Hindsight Audit: massiv überladen mit veralteter Infra-Topologie. Going-Forward-Policy: nur noch semantische Pointer + Entscheidungen
- LCM gesund: 7.590 messages, 34 DAG nodes, 21.2:1 compression ratio
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
## [2026-09-27] architecture | Laya Email-Organizer Migration + Session-Nutzung
- Rechnungen-Organizer Cron `f773f8c23230` von LLM-Agent → `no_agent` Script mit Laya migriert
- 12 Kategorien, 3-Schichten-Safety (Confidence-Gate + Subject-Validierung + Move-Erfolg)
- Globale Ordner (Rechnungen/Bestellungen/Gutschriften/Gutscheine) statt Monatssortierung
- Wiki-Seite `systems/laya.md` erstellt mit Usage Guide für Session-Nutzung
- User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score)
- Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md`
## [2026-09-17] workflow | CT108-Endausbau: Census, Ghost-Router CT99999, Alias-Repair, Stop
- **Census (auth):** K8s=25 / CT108=31 / gemeinsam=23 — 22 Tips identisch, dominik/memory=K8s-Superset (enthält CT-Tip de97cd6d), nur-CT=8× PoC-Müll. Archiv: 29 Bare-Bundles 143 MB unter `/home/debian/git-archive/ct108-final/` (2 Failures = legitim leere Repos).
- **Ghost-Router enttarnt:** CT99999 (Traefik-LXC auf proxmox7, 10.0.60.10) terminiert TLS für *.familie-schoen.com und forwardet plain HTTP an .203. `git.familie-schoen.com` → .105:3000 WAR der letzte Live-Konsument von CT108; `git.schoen.codes` lief längst per Double-Hop aufs K8s. Fix: gitea-service-Upstream → .203 (Backup `explicit-http.yml.bak-hermes-20260917`) + Alias-Host im K8s-Ingress (Commit `9e9b6ee`). Beide Hostnamen jetzt v1.27.0.
- **Stop vollzogen (genehmigt):** `pct stop 108` auf ms-a2-2 (10.0.20.93) — connection-refused-Beweis, Fleet 22/25 Synced, Runner unversehrt. Wiki-Lügen korrigiert (ip-map behauptete „stopped" seit Wochen).
- **Lessons:** 1P-SA braucht je Call `--vault`; `op read --reveal` existiert nicht (stdout=Secret); `op item get --reveal` maskiert nur Display (JSON-Captures intakt); `/repos/search` = `{ok,data}`-Envelope; blankes Token erzeugt glaubwürdig LEERE Census (Fast-Fehlentscheidung „K8s hat nur 2 Repos").
- Docs: `docs/solutions/workflows/2026-09-17-ct108-full-decommission-census-router-topology.md` · Wiki: systems/gitea.md, reference/ip-map.md
- Offen (je Freigabe): Zombie-Secret `argocd-repo-credentials` löschen; immich-Zwillings-App (toter rendered-Dump vom 01.08., Live gehört immich-config) entfernen.
## [2026-09-17] fix | Merge-Day-Kampagne abgeschlossen: Konvergenz + Autosync-Rennen + Live-State-Arbitrage
- **Konvergenz DONE:** Merge `d76d3e8` (fork 46fa168, 24.07.) + Fixups `4987341`/`8c741c5` auf BEIDEN Remotes (CT108 + K8s-Gitea SSH 10.0.30.200). Ahead37/behind63-Narrativ endgültig begraben (Cache-Phantom). Single-Remote-Ziel erreicht: beide Tipps identisch.
- **Autosync-Renne verarbeitet (4 Minentypen):** authelia Duplikat-Volume (union-merge) entfernt; homepage `authelia-auth`-Middleware restauriert (Jul-Entscheidung ging nie live); paperless plaintext-OIDC-Secret GELÖSCHT statt Wert-Rollback (ESO-Ownership seit 24.08., Quelle 1P); newborn-App ceph-csi-cephfs eingefroren (autosync-Block aus Git entfernt — Helm-Release 3.17.0 wartet auf Adoption; rbd-chart 3.10.1 weiterhin ungoverned, Backlog).
- **3 Sync-Failures via Live-State-Arbitrage geheilt (alle Synced/Healthy @ 8c741c5):** (1) backups: VSC-CRD-Flavor-Falle — RKE2-Addon-CRD ist FLAT (driver/deletionPolicy top-level, `spec:` verboten!), Jul-Files waren für Upstream-Flavor korrekt; beide Files geflattet, cephfs-snapclass erstmals LIVE ERZEUGT. (2) databases: CNPG grow-only — Git auf 100Gi/20Gi hochaligned (Shrink verboten). (3) hindsight: VCT storageClassName immutable (hdd→flash aligniert), Live-Hotfixes gespiegelt (Node-Pin cp-01, ReadinessProbe draußen, secretKeyRef statt Klartext-PW). hindsight-postgres-0 rotierte sauber, PVC 47d Bound, API pollt.
- **Push-Learned:** canonical-Push braucht EXPLIZITEN Key (`GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o IdentitiesOnly=yes"`) — kein ~/.ssh/config vorhanden, Default-Key lehnt ab.
- **Chronisch (prä-merge, offen):** kube-prometheus-stack Synced/FAILED (CRD-Annotation >256kB); residual OutOfSync gitea-config (Deployment/gitea-runner) + immich-config (SA/CM-Reste) = Alt-Backlog, keine Regression.
- **Finale Ordnung (nächste Schritte):** ① ArgoCD-Sources auf ssh://git@10.0.30.200:22 flippen (incl. Secret argocd-repo-credentials) ② CT108 final stoppen (explizite Freigabe Dominik) ③ CSI-Adoption ④ kpstack-CRD-Fix.
- Docs: `docs/solutions/workflows/2026-09-17-argocd-autosync-race-four-mine-types.md` + `docs/solutions/bug-fixes/2026-09-17-live-state-arbitration-crd-flavors-hotfix-mirroring.md` · Morgen-Doc (.202-Empfehlung) korrigiert auf .200.
## [2026-09-17] bug-fix | Ghost-Instance-Regression: CT108-Gitea als ArgoCD-Source + gitea-backup RCA-Quality
- gitea-backup-29826930 (03:30Z) failed: BackoffLimitExceeded, 3 Instant-Crashes in 73s. Zwei Auto-RCAs attribuierten auf 1P/ESO-Rate-Limit — STRUKTURELL widerlegt (crashing init-container gitea-files konsumiert keine Secrets; mounted Secrets intakt).
- Echter Fund: ArgoCD-Apps (root/proxy/paperless/schoenkitchen/gitea-config) zogen von http://10.0.30.105:3000 = CT108-Ghost (nach Dekommissionierung 24.07. wieder eingeschaltet, stale Mirror ohne Hardening-Commit 2014673) → "Synced" maskierte fehlende failedJobsHistoryLimit-Felder live.
- Sofortmassnahme: manueller Rerun gitea-backup-manual-161538 SUCCESS (201,7 MB, S3-Upload verifiziert), Failed-Job gelöscht, Alert clear.
- Offen (awaiting owner): ArgoCD-Sources auf ssh://git@10.0.30.202:22 flippen, Repo-Divergenz ahead37/behind63 + tote HTTPS-Auth forensisch, CT108 final stoppen, ESO-Nachtsättigung (00:00Z-Fenster, cf. PAT-003) analysieren.
- Doc: `docs/solutions/bug-fixes/2026-09-17-gitops-ghost-instance-regression-masked-hardening.md` · Wiki: systems/gitea (CT108-Status), concepts/gitops-workflow (Single-Remote-Regel)
- **Abend-Phase:** SSH-Deny-Rootcause = keine Keys/Tokens im K8s-Gitea-DB registriert (Opfer der 01.08.-Migration) → via 1P-Token (hermes-gitops) re-registriert: User-Key + ArgoCD-Deploy-Key (beide end-to-end verifiziert, push dry-run OK). Echter Branch-Split am 24.07. entdeckt: K8s-Gitea-main=29.07.-Stand (b9d7441, 70 Commits incl. mariadb:11.4-Fix für gitea-backup), CT108/local=38 Commits ab 04.09. Hardening TTL=86400 + Exit-42-Guard via ArgoCD gelanded (a0c9b6b). Heutiger DB-Dump validiert (116 Tables, kompletter Trailer). Nächste Schritte: Historien-Konvergenz → Source-Flip → CT108-Stop (Freigabe).
## [2026-08-30] retro | Compound Learning Retrospective (Last 30 Days)
- Reviewed sessions from Jul 31 – Aug 30, 2026
- **4 new solution docs** written by subagents:
- `bug-fixes/2026-08-05-ceph-squid-ec-pool-mark-complete-bug.md` — EC pool mark-complete fails in Squid
- `bug-fixes/2026-08-06-traefik-cross-namespace-routing-pitfall.md` — Two Traefik 404 incidents
- `architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md` — Full RKE2 rebuild sequence
- `architecture/2026-08-01-immich-data-loss-no-backup.md` — Ceph redundancy ≠ backup
- **2 new patterns** extracted:
- PAT-008: Ansible default_ipv6 fact missing on fresh VMs
- PAT-009: Stale NBD devices after RBD volume swap
- **4 Hindsight entries** indexed with solution summaries
- Key themes: K8s cluster disaster recovery, Ceph EC pool unrecoverability, Traefik routing complexity, backup gap identification
## [2026-08-30] feat | WikiSkill Patterns Directory + Skill-Impact Tracker
- Analysed arXiv:2608.27454 (WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution)
- **New `patterns/` directory** in LLM Wiki — structured failure-mode patterns inspired by WikiSkill's Wiki Layer
- Created 7 initial pattern pages from existing MEMORY.md entries + solution docs:
- PAT-001: Galera DDL TOI Deadlock
- PAT-002: Traefik reload unreliable
- PAT-003: 1Password account-level rate-limit
- PAT-004: VFIO GPU passthrough race condition
- PAT-005: Finanzblick sync WAF-blocked
- PAT-006: LinkedIn React nativeInputSetter
- PAT-007: Ceph EC pool + SSD wear-level
- **Skill-Impact Tracker** (`patterns/skill-impact.md`) — audit trail for skill modifications (accept/reject history)
- **Template** (`patterns/_template.md`) for future pattern creation
- **compound-learning skill patched** — added Phase 4.8 (Pattern Extraction) + Phase 4.9 (Skill-Impact Tracking)
- Updated index.md with Patterns section
- Hindsight: indexed with tags architecture, wikiskill, patterns, skill-evolution
## [2026-08-13] feat | Memory Architecture Enhancement (arxiv 2602.06052v4)
- Implemented 3 new automated memory capabilities from survey paper analysis:
- **Decay & Importance Scoring** (`hindsight_decay_scoring.py`): Monthly cron (1st 03:30). Score = log(1+access)/log(51) × 0.5^(age/90d). Archives score<0.15 with 0 accesses. First run: 200 evaluated, 50 archived.
- **User Drift Detection** (`user_drift_detection.py`): Weekly cron (Mon 04:00). Compares 14d vs 90d baseline for message length, frustration, corrections, topic shifts, style changes. First run: detected 3 drifts (41% longer msgs, new focus gitops, declining ai_ml).
- **Skill Health Check** (`skill_health_check.py`): Weekly cron (Mon 04:30). Scans 187 active SKILL.md files for stubs, broken refs, missing pitfalls, staleness, duplicates. First run: 70 healthy, 117 warnings, 0 critical.
- Cron jobs: 112c106c0954, 280458a2b035, 761b6694d99b (all no_agent, deliver=origin)
- Wiki updated: systems/hindsight.md (meta-memory automation table expanded)
- Hindsight: indexed with tags architecture, memory-system, cron, arxiv-2602.06052
## [2026-08-13] fix | Seafile Recovery + Paperless Upgrade + Authelia Ingress
- **Gitea SSH access**: No SSH key on Hermes VM was authorized for Gitea. HTTP port 3000 has no external IngressRoute. Generated dedicated ED25519 key (`id_ed25519_gitea-hermes-push`), stored in 1Password (ID: lqaebdtewj5xatvpzfrkag73v4), registered in Gitea as key ID 4 for user dominik. SSH via `10.0.30.202:22` (LoadBalancer) is the reliable push path.
- **Paperless v3 IngressRoute**: Host was `dokumente.familie-schoen.com` instead of `dokumente-neu.familie-schoen.com`. Two-track fix: kubectl patch (immediate) + Git commit `fa4184b` (permanent via ArgoCD self-heal). Backtick escaping in Traefik Host() match requires `--patch-file` not inline `--patch`.
- Wiki updated: systems/gitea.md (SSH push key section), reference/ssh-keys.md (new key entry)
- Solution docs: workflows/2026-07-26-gitea-ssh-access-for-hermes.md, bug-fixes/2026-07-26-paperless-ingressroute-hostname-fix.md
- Hindsight: both solutions indexed with tags
## [2026-07-25] fix | Loki 500 Error + CNPG Leader Election + Stale Pods Cleanup
- **Loki HTTP 500**: `replication_factor: 3` in hash ring with only 1 SingleBinary instance → "too many unhealthy instances in the ring". Chart v6.42.0 ignores `loki.common.replication_factor` — correct path is `loki.commonConfig.replication_factor`. Fix: commit `cb70d36`
- **Memcached caches disabled**: `chunksCache.enabled: false`, `resultsCache.enabled: false` (SingleBinary doesn't need them). Commit `f8852e8`
- **Fluent Bit**: cascading failure from Loki 500s — fixed automatically once Loki accepted pushes
- **CNPG 28+129 restarts**: Leader election lease renewal failed during API server latency spikes (Ceph recovery I/O). Default 15s/10s too short. Fix: `--leader-lease-duration=60 --leader-renew-deadline=40` via `additionalArgs`. Commit `6ed4a53`
- **17 stale node-debugger pods** deleted from default namespace
- **OSDs 0+2 destroyed+purged** (proxmox2, SSDs with 92% wear + slow ops). CRUSH host proxmox2 removed. 13 OSDs remaining on 8 hosts.
- Wiki updated: ceph-cluster.md, loki-fluentbit.md
## [2026-07-25] fix | Memory Sync Broken — HTTP→HTTPS + Ceph Duplicate
- Git push failed: remote URL used `http://` but Gitea redirects to `https://` — git doesn't follow auth redirects
- Token itself was valid (same as ArgoCD `argocd-repo-credentials`), just wrong protocol
- Fixed: `git remote set-url` to `https://`
- Eliminated duplicate: `systems/ceph.md` (82 lines, no frontmatter) merged into `systems/ceph-cluster.md` (canonical, referenced in index.md)
- Fixed sync script: removed `--allow-empty` flag, improved change detection
## [2026-07-24] fix | Velero CSI Volume Snapshots Three-Layer Fix
- Velero backups PartiallyFailed for 121 days — all PVCs skipped
- Fix 1: Created VolumeSnapshotClass `ceph-rbd-snapclass` (rbd.csi.ceph.com)
- Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22)
- Gitea SSH user is `git` not `gitea`; git.schoen.codes → Traefik (10.0.30.208), SSH on 10.0.30.202
- New SSH deploy key generated, added to Gitea repo
- Fix 3: `features: EnableCSI` must be under `configuration:` in Velero Helm values (not top-level)
- Must be string, not array — array form breaks Helm template
- Verified: test backup created VolumeSnapshot + VolumeSnapshotContent, uploaded to S3 ✅
- Solution doc: docs/solutions/bug-fixes/2026-07-24-velero-csi-snapshots-three-layer-fix.md
- Commits: 5ebfa30, 19f35bc, cd18c58, baf85f5, f652bfb
## [2026-07-24] fix | VFIO Module Loading Race Condition RCA + Fix
- RCA: GPU passthrough failed on ms-a2-1 due to missing modprobe config (NOT kernel version)
- Root cause: Missing `softdep amdgpu pre: vfio-pci` + incomplete DRM blacklist → race condition (302s vs 2.4s bind time)
- Fix applied to ms-a2-1: Added `blacklist drm`, `blacklist drm_kms_helper`, `softdep amdgpu pre: vfio-pci`, rebuilt initramfs
- Verification pending (reboot required, 4 VMs + OSD 13 on ms-a2-1)
- Solution doc: docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md
- Also documented: K8s persistent storage architecture (all PVCs on Ceph RBD, 3x replication)
- Solution doc: docs/solutions/architecture/2026-07-24-k8s-persistent-storage-ceph-rbd.md
## [2026-07-24] fix | worker-05 GPU Passthrough — Move to ms-a2-2
- Problem: worker-05 (VM 102) on ms-a2-1 had kernel BUG with 1002:13c0 GPU — /dev/dri/renderD128 missing
- ms-a2-1 and ms-a2-2 have identical hardware (AMD Ryzen 9 9955HX, GPU 1002:13c0)
- Solution: Moved VM 102 from ms-a2-1 to ms-a2-2 (offline migration, Ceph RBD shared storage)
- Result: renderD128 ✅ visible in guest, amdgpu driver loaded, all pods running
- IaC: main.tf cleaned up — single unified rke2_worker block, worker-05 node=ms-a2-2, worker-06 removed
- PCI mapping: amd-gpu now includes n5pro + ms-a2-1 + ms-a2-2
- Commit: 21bd335
- Also: ms-a2-2 upgraded PVE 9.1.1→9.2.5, kernel 7.0.14-6-pve, VFIO persistent after reboot, HA integrated
## [2026-07-24] init | ms-a2-2 Node Initialization
- New Proxmox node ms-a2-2 (10.0.20.93, nodeid 1) joined cluster
- Hardware: AMD Ryzen 9 9955HX (32 threads), 91GB RAM, 1.8TB NVMe
- GPU: AMD Radeon 1002:13c0 (identical to ms-a2-1) — VFIO passthrough configured
- PCI mapping: amd-gpu-ms-a2-2 created, added to generic amd-gpu mapping
- Ceph OSD 14: created on nvme0n1 (1.8TB SSD), class ssd, weight 1.82
- Fluent-bit: installed + configured (hostname adapted to ms-a2-2)
- APT sources: fixed from enterprise to no-subscription (ceph + pve)
- LLM Wiki: ip-map.md, ceph.md updated; AGENTS.md updated (9 nodes)
- Note: PVE 9.1.1 on ms-a2-2 (rest of cluster is 9.2.3) — upgrade pending
## [2026-07-24] cleanup | proxmox1 Permanent Removal + Ceph PG Fixes
- Removed proxmox1 from Ceph CRUSH (empty host bucket, weight 0)
- Also removed stale empty buckets: proxmox, px-tmp20
- Updated corosync.conf: removed proxmox1 node entry, config_version 14→15, expected_votes 9→8
- Removed stale /etc/pve/nodes/{proxmox1,proxmox,pve,px-tmp20}
- Updated IaC: proxmox_api_url 10.0.20.10→10.0.20.91, template_node proxmox1→n5pro, target_nodes list updated
- Updated AGENTS.md topology (9→8 nodes), LLM Wiki ip-map.md + ceph.md
- Commit: 469d29d
- Additional: Manual pg-upmap for stuck PGs 5.13→[1,6,7], 6.6c→[11,8,7] — all clean+remapped eliminated
## [2026-07-24] fix | Ceph Remapped PGs — Weight Imbalance + pg_num Mismatch
- RCA: 11 active+clean+remapped PGs caused by CRUSH placement failure
- Root cause 1: Pool 6 (hdd_disk) pg_num=120 ≠ pgp_num=112 (autoscaler mid-merge)
- Root cause 2: Extreme HDD host weight imbalance (n5pro=10TB vs proxmox6=0.3TB)
- Fix: Equalized pg_num→112, set nopgchange=true, reweighted osd.7 0.80→1.0
- Result: 9/11 PGs recovering (backfilling at 36MiB/s), 2 cosmetic clean+remapped remain
- Created systems/ceph.md with full OSD/pool/CRUSH documentation
## [2026-07-24] update | Gitea Actions K8s Runner Bug + DinD Pitfalls
- Documented Gitea 1.27.0 bug: actions_ready_job queue doesn't assign tasks to K8s-based runners
- Updated systems/gitea.md with known issue, DinD pitfalls, runner management via DB
- Solution doc: docs/solutions/bug-fixes/2026-07-24-gitea-actions-k8s-runner-task-assignment-bug.md
- IaC committed: runner-deployment.yaml + runner-config.yaml (DinD TCP sidecar pattern)
## [2026-04-28] init | Memory System Created
- Created memory directory structure at ~/.hermes/memory
- Installed qmd 2.1.0 (Query Markup Documents)
- Created initial index.md with entities, concepts, and sources
- Created Entity pages: Infrastructure, Email-System, Health-Fitness, Team-Structure, Memory-System
- Created Concept pages: Network-Architecture, Email-Organization, Monitoring-System, Proxmox-Cluster, qmd-Knowledge-Base
- Initialized git repo, pushed to origin/main
## [2026-04-29] fix | Memory Process Established
- Corrected stale memory-entry-001.md
- Established protocol: memory update + git sync on every session
## [2026-04-30] ingest | cloud.familie-schoen.com
- Full port scan, DNS, SSL, HTTP headers, API analysis
- Identified: Seafile 13.0.19 on nginx, Let's Encrypt SSL
## [2026-07-24] restructure | Full Wiki Restructuring
- Deleted 5 root-duplicate files (dominik-schoen, email-organization, team-structure, llm-wiki-pattern, embedding-model)
- Deleted obsolete Entities/Memory-System.md, Concepts/qmd-Knowledge-Base.md
- Deleted raw/articles/ and memory-entry-001/002.md (stale)
- Renamed Entities/ → entities/, Concepts/ → concepts/ (lowercase)
- Created new directories: systems/, reference/
- Moved Proxmox-Cluster, Monitoring-System, seafile to systems/
- Updated all existing pages with current data (April → July 2026):
- Infrastructure: 2 Nodes → 9 Nodes, added Ceph/RKE2/Galera/Loki
- Proxmox-Cluster: PVE 9.2.3, 9 Nodes, Fluent Bit details, PVE Tasks, Ceph Audit
- Network-Architecture: cleaned up, links to reference pages
- Monitoring: Prometheus/Grafana/Loki/HolmesGPT (was just "weekly network scans")
- Created 6 new system pages: ceph-cluster, rke2-kubernetes, galera-maxscale, loki-fluentbit, gitea, hindsight
- Created 3 reference pages: ip-map, ssh-keys, ports
- Created 2 concept pages: gitops-workflow, credential-policy
- Rewrote index.md with new structure + Memory Layer Architecture table
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
## [2026-09-29] update | VM302 Zombie-Recovery nach Migration
- VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot
- Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause
- Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min
- Endstand: 3/3 Synced, Primary, MaxScale alle Server Running
- Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md
- Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)