# Memory Log ## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur - Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen. - Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate. - SMART-Korrektur: Kingston SFYRDK2000G (osd.2-Träger) ist GESUND (wear 6 %, media_errors 0, spare 100 %, PoH 1469) — frühere "Austausch-Kandidatin"-These RETRAKIERT. Vorfall war Controller-/Fabric-Ebene. Neu beobachten: Temps 70/78 °C, T1-throttle 3×. - PAT-014 erweitert (Smart-Log-Abschnitt), ceph-cluster.md korrigiert, Commit cc2f237. ## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014) - osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin. *(Korrektur später am selben Tag: SMART-Audit → gesund, kein Austausch; siehe nächster Eintrag.)* - osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend. - ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben. - Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014. ## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013) - Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün. - Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok. - **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate. - Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit . ## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012) - Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok. - Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom). - Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172. ## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011) - Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON). - Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200. - Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body. - Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales). ## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen - proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%. - Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets. - VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP. - Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt. ## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet - GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md. - Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc. - proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen. - Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert. ## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010) - Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h - Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB) - Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX - Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared - Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr - Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return - Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010) ## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live - VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil) - hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key) - E2E verifiziert: Login 200, Agent-Antwort via vllm/release/glm-5-2 @ ai.noris.de - PITFALL: hermes-agent pyproject pinnt Core-Deps hinter python_version>='3.14'-Markers → auf Python 3.12 installiert `pip install -e .` NULL Deps ("AIAgent not available"). Fix: manuelle Pin-Installation + Missing-Import-Loop bis `import run_agent` - Backup: täglich 23:00 Job backup-2e8a34e3-66cb (all=1) deckt VM107; manueller Verify TASK OK 48s - Wiki: systems/sarah-hermes.md, ip-map.md erweitert; iac-homelab commits 0d5f017/1ce4816 - OFFEN: Telegram-Bot blockiert auf BotFather-Token (Sarah/Dominik) ## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret - Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params - CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx - Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min) - Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05 - Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2) - Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes) - Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md ## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit - Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash - Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true` - Gitea Pod + schoenkitchen runner durch Pod-Delete wiederhergestellt - ArgoCD Audit: 17/25 Apps via Gitea SSH, 20/25 Synced+Healthy - Known Issues: `gitea`/`authelia` Health=Progressing (Ingress-LB-IP Gap), `immich` doppelt verwaltet, `gitea-config` Deployment Drift - `DEFAULT_ACTIONS_URL=https://gitea.com` deprecated → Helm override auf `self` empfohlen - Wiki `systems/gitea.md` aktualisiert: Runner Status, Incident, ArgoCD Issues - Solution Doc: `docs/solutions/bug-fixes/2026-09-27-gitea-csi-rbac-volumeattachment-fix.md` ## [2026-09-27] architecture | Memory Layer Restructuring + Laya Session-Nutzung - USER.md bereinigt: Infra-Fakten entfernt, nur noch User-Preferences/Safety-Rules (1.087/1.375 chars) - MEMORY.md ausgedünnt: 2.145→1.664 chars, alle Infra-Details zeigen auf Wiki-Seiten mit `→Wiki` Pointern - 4 neue Wiki-Seiten: systems/noris-ai, systems/frigate, systems/homeassistant, systems/paperless - Neues Concept: concepts/memory-layer-architecture — Entscheidungsbaum "was wohin gehört" - Hindsight Audit: massiv überladen mit veralteter Infra-Topologie. Going-Forward-Policy: nur noch semantische Pointer + Entscheidungen - LCM gesund: 7.590 messages, 34 DAG nodes, 21.2:1 compression ratio - User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score) ## [2026-09-27] architecture | Laya Email-Organizer Migration + Session-Nutzung - Rechnungen-Organizer Cron `f773f8c23230` von LLM-Agent → `no_agent` Script mit Laya migriert - 12 Kategorien, 3-Schichten-Safety (Confidence-Gate + Subject-Validierung + Move-Erfolg) - Globale Ordner (Rechnungen/Bestellungen/Gutschriften/Gutscheine) statt Monatssortierung - Wiki-Seite `systems/laya.md` erstellt mit Usage Guide für Session-Nutzung - User-Preference: Laya künftig in normalen Sessions nutzen (choice/noul/score) - Solution Doc: `docs/solutions/architecture/2026-09-27-laya-email-organizer-migration.md` ## [2026-09-17] workflow | CT108-Endausbau: Census, Ghost-Router CT99999, Alias-Repair, Stop - **Census (auth):** K8s=25 / CT108=31 / gemeinsam=23 — 22 Tips identisch, dominik/memory=K8s-Superset (enthält CT-Tip de97cd6d), nur-CT=8× PoC-Müll. Archiv: 29 Bare-Bundles 143 MB unter `/home/debian/git-archive/ct108-final/` (2 Failures = legitim leere Repos). - **Ghost-Router enttarnt:** CT99999 (Traefik-LXC auf proxmox7, 10.0.60.10) terminiert TLS für *.familie-schoen.com und forwardet plain HTTP an .203. `git.familie-schoen.com` → .105:3000 WAR der letzte Live-Konsument von CT108; `git.schoen.codes` lief längst per Double-Hop aufs K8s. Fix: gitea-service-Upstream → .203 (Backup `explicit-http.yml.bak-hermes-20260917`) + Alias-Host im K8s-Ingress (Commit `9e9b6ee`). Beide Hostnamen jetzt v1.27.0. - **Stop vollzogen (genehmigt):** `pct stop 108` auf ms-a2-2 (10.0.20.93) — connection-refused-Beweis, Fleet 22/25 Synced, Runner unversehrt. Wiki-Lügen korrigiert (ip-map behauptete „stopped" seit Wochen). - **Lessons:** 1P-SA braucht je Call `--vault`; `op read --reveal` existiert nicht (stdout=Secret); `op item get --reveal` maskiert nur Display (JSON-Captures intakt); `/repos/search` = `{ok,data}`-Envelope; blankes Token erzeugt glaubwürdig LEERE Census (Fast-Fehlentscheidung „K8s hat nur 2 Repos"). - Docs: `docs/solutions/workflows/2026-09-17-ct108-full-decommission-census-router-topology.md` · Wiki: systems/gitea.md, reference/ip-map.md - Offen (je Freigabe): Zombie-Secret `argocd-repo-credentials` löschen; immich-Zwillings-App (toter rendered-Dump vom 01.08., Live gehört immich-config) entfernen. ## [2026-09-17] fix | Merge-Day-Kampagne abgeschlossen: Konvergenz + Autosync-Rennen + Live-State-Arbitrage - **Konvergenz DONE:** Merge `d76d3e8` (fork 46fa168, 24.07.) + Fixups `4987341`/`8c741c5` auf BEIDEN Remotes (CT108 + K8s-Gitea SSH 10.0.30.200). Ahead37/behind63-Narrativ endgültig begraben (Cache-Phantom). Single-Remote-Ziel erreicht: beide Tipps identisch. - **Autosync-Renne verarbeitet (4 Minentypen):** authelia Duplikat-Volume (union-merge) entfernt; homepage `authelia-auth`-Middleware restauriert (Jul-Entscheidung ging nie live); paperless plaintext-OIDC-Secret GELÖSCHT statt Wert-Rollback (ESO-Ownership seit 24.08., Quelle 1P); newborn-App ceph-csi-cephfs eingefroren (autosync-Block aus Git entfernt — Helm-Release 3.17.0 wartet auf Adoption; rbd-chart 3.10.1 weiterhin ungoverned, Backlog). - **3 Sync-Failures via Live-State-Arbitrage geheilt (alle Synced/Healthy @ 8c741c5):** (1) backups: VSC-CRD-Flavor-Falle — RKE2-Addon-CRD ist FLAT (driver/deletionPolicy top-level, `spec:` verboten!), Jul-Files waren für Upstream-Flavor korrekt; beide Files geflattet, cephfs-snapclass erstmals LIVE ERZEUGT. (2) databases: CNPG grow-only — Git auf 100Gi/20Gi hochaligned (Shrink verboten). (3) hindsight: VCT storageClassName immutable (hdd→flash aligniert), Live-Hotfixes gespiegelt (Node-Pin cp-01, ReadinessProbe draußen, secretKeyRef statt Klartext-PW). hindsight-postgres-0 rotierte sauber, PVC 47d Bound, API pollt. - **Push-Learned:** canonical-Push braucht EXPLIZITEN Key (`GIT_SSH_COMMAND="ssh -i ~/.ssh/id_ed25519_gitea-hermes-push -o IdentitiesOnly=yes"`) — kein ~/.ssh/config vorhanden, Default-Key lehnt ab. - **Chronisch (prä-merge, offen):** kube-prometheus-stack Synced/FAILED (CRD-Annotation >256kB); residual OutOfSync gitea-config (Deployment/gitea-runner) + immich-config (SA/CM-Reste) = Alt-Backlog, keine Regression. - **Finale Ordnung (nächste Schritte):** ① ArgoCD-Sources auf ssh://git@10.0.30.200:22 flippen (incl. Secret argocd-repo-credentials) ② CT108 final stoppen (explizite Freigabe Dominik) ③ CSI-Adoption ④ kpstack-CRD-Fix. - Docs: `docs/solutions/workflows/2026-09-17-argocd-autosync-race-four-mine-types.md` + `docs/solutions/bug-fixes/2026-09-17-live-state-arbitration-crd-flavors-hotfix-mirroring.md` · Morgen-Doc (.202-Empfehlung) korrigiert auf .200. ## [2026-09-17] bug-fix | Ghost-Instance-Regression: CT108-Gitea als ArgoCD-Source + gitea-backup RCA-Quality - gitea-backup-29826930 (03:30Z) failed: BackoffLimitExceeded, 3 Instant-Crashes in 73s. Zwei Auto-RCAs attribuierten auf 1P/ESO-Rate-Limit — STRUKTURELL widerlegt (crashing init-container gitea-files konsumiert keine Secrets; mounted Secrets intakt). - Echter Fund: ArgoCD-Apps (root/proxy/paperless/schoenkitchen/gitea-config) zogen von http://10.0.30.105:3000 = CT108-Ghost (nach Dekommissionierung 24.07. wieder eingeschaltet, stale Mirror ohne Hardening-Commit 2014673) → "Synced" maskierte fehlende failedJobsHistoryLimit-Felder live. - Sofortmassnahme: manueller Rerun gitea-backup-manual-161538 SUCCESS (201,7 MB, S3-Upload verifiziert), Failed-Job gelöscht, Alert clear. - Offen (awaiting owner): ArgoCD-Sources auf ssh://git@10.0.30.202:22 flippen, Repo-Divergenz ahead37/behind63 + tote HTTPS-Auth forensisch, CT108 final stoppen, ESO-Nachtsättigung (00:00Z-Fenster, cf. PAT-003) analysieren. - Doc: `docs/solutions/bug-fixes/2026-09-17-gitops-ghost-instance-regression-masked-hardening.md` · Wiki: systems/gitea (CT108-Status), concepts/gitops-workflow (Single-Remote-Regel) - **Abend-Phase:** SSH-Deny-Rootcause = keine Keys/Tokens im K8s-Gitea-DB registriert (Opfer der 01.08.-Migration) → via 1P-Token (hermes-gitops) re-registriert: User-Key + ArgoCD-Deploy-Key (beide end-to-end verifiziert, push dry-run OK). Echter Branch-Split am 24.07. entdeckt: K8s-Gitea-main=29.07.-Stand (b9d7441, 70 Commits incl. mariadb:11.4-Fix für gitea-backup), CT108/local=38 Commits ab 04.09. Hardening TTL=86400 + Exit-42-Guard via ArgoCD gelanded (a0c9b6b). Heutiger DB-Dump validiert (116 Tables, kompletter Trailer). Nächste Schritte: Historien-Konvergenz → Source-Flip → CT108-Stop (Freigabe). ## [2026-08-30] retro | Compound Learning Retrospective (Last 30 Days) - Reviewed sessions from Jul 31 – Aug 30, 2026 - **4 new solution docs** written by subagents: - `bug-fixes/2026-08-05-ceph-squid-ec-pool-mark-complete-bug.md` — EC pool mark-complete fails in Squid - `bug-fixes/2026-08-06-traefik-cross-namespace-routing-pitfall.md` — Two Traefik 404 incidents - `architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md` — Full RKE2 rebuild sequence - `architecture/2026-08-01-immich-data-loss-no-backup.md` — Ceph redundancy ≠ backup - **2 new patterns** extracted: - PAT-008: Ansible default_ipv6 fact missing on fresh VMs - PAT-009: Stale NBD devices after RBD volume swap - **4 Hindsight entries** indexed with solution summaries - Key themes: K8s cluster disaster recovery, Ceph EC pool unrecoverability, Traefik routing complexity, backup gap identification ## [2026-08-30] feat | WikiSkill Patterns Directory + Skill-Impact Tracker - Analysed arXiv:2608.27454 (WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution) - **New `patterns/` directory** in LLM Wiki — structured failure-mode patterns inspired by WikiSkill's Wiki Layer - Created 7 initial pattern pages from existing MEMORY.md entries + solution docs: - PAT-001: Galera DDL TOI Deadlock - PAT-002: Traefik reload unreliable - PAT-003: 1Password account-level rate-limit - PAT-004: VFIO GPU passthrough race condition - PAT-005: Finanzblick sync WAF-blocked - PAT-006: LinkedIn React nativeInputSetter - PAT-007: Ceph EC pool + SSD wear-level - **Skill-Impact Tracker** (`patterns/skill-impact.md`) — audit trail for skill modifications (accept/reject history) - **Template** (`patterns/_template.md`) for future pattern creation - **compound-learning skill patched** — added Phase 4.8 (Pattern Extraction) + Phase 4.9 (Skill-Impact Tracking) - Updated index.md with Patterns section - Hindsight: indexed with tags architecture, wikiskill, patterns, skill-evolution ## [2026-08-13] feat | Memory Architecture Enhancement (arxiv 2602.06052v4) - Implemented 3 new automated memory capabilities from survey paper analysis: - **Decay & Importance Scoring** (`hindsight_decay_scoring.py`): Monthly cron (1st 03:30). Score = log(1+access)/log(51) × 0.5^(age/90d). Archives score<0.15 with 0 accesses. First run: 200 evaluated, 50 archived. - **User Drift Detection** (`user_drift_detection.py`): Weekly cron (Mon 04:00). Compares 14d vs 90d baseline for message length, frustration, corrections, topic shifts, style changes. First run: detected 3 drifts (41% longer msgs, new focus gitops, declining ai_ml). - **Skill Health Check** (`skill_health_check.py`): Weekly cron (Mon 04:30). Scans 187 active SKILL.md files for stubs, broken refs, missing pitfalls, staleness, duplicates. First run: 70 healthy, 117 warnings, 0 critical. - Cron jobs: 112c106c0954, 280458a2b035, 761b6694d99b (all no_agent, deliver=origin) - Wiki updated: systems/hindsight.md (meta-memory automation table expanded) - Hindsight: indexed with tags architecture, memory-system, cron, arxiv-2602.06052 ## [2026-08-13] fix | Seafile Recovery + Paperless Upgrade + Authelia Ingress - **Gitea SSH access**: No SSH key on Hermes VM was authorized for Gitea. HTTP port 3000 has no external IngressRoute. Generated dedicated ED25519 key (`id_ed25519_gitea-hermes-push`), stored in 1Password (ID: lqaebdtewj5xatvpzfrkag73v4), registered in Gitea as key ID 4 for user dominik. SSH via `10.0.30.202:22` (LoadBalancer) is the reliable push path. - **Paperless v3 IngressRoute**: Host was `dokumente.familie-schoen.com` instead of `dokumente-neu.familie-schoen.com`. Two-track fix: kubectl patch (immediate) + Git commit `fa4184b` (permanent via ArgoCD self-heal). Backtick escaping in Traefik Host() match requires `--patch-file` not inline `--patch`. - Wiki updated: systems/gitea.md (SSH push key section), reference/ssh-keys.md (new key entry) - Solution docs: workflows/2026-07-26-gitea-ssh-access-for-hermes.md, bug-fixes/2026-07-26-paperless-ingressroute-hostname-fix.md - Hindsight: both solutions indexed with tags ## [2026-07-25] fix | Loki 500 Error + CNPG Leader Election + Stale Pods Cleanup - **Loki HTTP 500**: `replication_factor: 3` in hash ring with only 1 SingleBinary instance → "too many unhealthy instances in the ring". Chart v6.42.0 ignores `loki.common.replication_factor` — correct path is `loki.commonConfig.replication_factor`. Fix: commit `cb70d36` - **Memcached caches disabled**: `chunksCache.enabled: false`, `resultsCache.enabled: false` (SingleBinary doesn't need them). Commit `f8852e8` - **Fluent Bit**: cascading failure from Loki 500s — fixed automatically once Loki accepted pushes - **CNPG 28+129 restarts**: Leader election lease renewal failed during API server latency spikes (Ceph recovery I/O). Default 15s/10s too short. Fix: `--leader-lease-duration=60 --leader-renew-deadline=40` via `additionalArgs`. Commit `6ed4a53` - **17 stale node-debugger pods** deleted from default namespace - **OSDs 0+2 destroyed+purged** (proxmox2, SSDs with 92% wear + slow ops). CRUSH host proxmox2 removed. 13 OSDs remaining on 8 hosts. - Wiki updated: ceph-cluster.md, loki-fluentbit.md ## [2026-07-25] fix | Memory Sync Broken — HTTP→HTTPS + Ceph Duplicate - Git push failed: remote URL used `http://` but Gitea redirects to `https://` — git doesn't follow auth redirects - Token itself was valid (same as ArgoCD `argocd-repo-credentials`), just wrong protocol - Fixed: `git remote set-url` to `https://` - Eliminated duplicate: `systems/ceph.md` (82 lines, no frontmatter) merged into `systems/ceph-cluster.md` (canonical, referenced in index.md) - Fixed sync script: removed `--allow-empty` flag, improved change detection ## [2026-07-24] fix | Velero CSI Volume Snapshots Three-Layer Fix - Velero backups PartiallyFailed for 121 days — all PVCs skipped - Fix 1: Created VolumeSnapshotClass `ceph-rbd-snapclass` (rbd.csi.ceph.com) - Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22) - Gitea SSH user is `git` not `gitea`; git.schoen.codes → Traefik (10.0.30.208), SSH on 10.0.30.202 - New SSH deploy key generated, added to Gitea repo - Fix 3: `features: EnableCSI` must be under `configuration:` in Velero Helm values (not top-level) - Must be string, not array — array form breaks Helm template - Verified: test backup created VolumeSnapshot + VolumeSnapshotContent, uploaded to S3 ✅ - Solution doc: docs/solutions/bug-fixes/2026-07-24-velero-csi-snapshots-three-layer-fix.md - Commits: 5ebfa30, 19f35bc, cd18c58, baf85f5, f652bfb ## [2026-07-24] fix | VFIO Module Loading Race Condition RCA + Fix - RCA: GPU passthrough failed on ms-a2-1 due to missing modprobe config (NOT kernel version) - Root cause: Missing `softdep amdgpu pre: vfio-pci` + incomplete DRM blacklist → race condition (302s vs 2.4s bind time) - Fix applied to ms-a2-1: Added `blacklist drm`, `blacklist drm_kms_helper`, `softdep amdgpu pre: vfio-pci`, rebuilt initramfs - Verification pending (reboot required, 4 VMs + OSD 13 on ms-a2-1) - Solution doc: docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md - Also documented: K8s persistent storage architecture (all PVCs on Ceph RBD, 3x replication) - Solution doc: docs/solutions/architecture/2026-07-24-k8s-persistent-storage-ceph-rbd.md ## [2026-07-24] fix | worker-05 GPU Passthrough — Move to ms-a2-2 - Problem: worker-05 (VM 102) on ms-a2-1 had kernel BUG with 1002:13c0 GPU — /dev/dri/renderD128 missing - ms-a2-1 and ms-a2-2 have identical hardware (AMD Ryzen 9 9955HX, GPU 1002:13c0) - Solution: Moved VM 102 from ms-a2-1 to ms-a2-2 (offline migration, Ceph RBD shared storage) - Result: renderD128 ✅ visible in guest, amdgpu driver loaded, all pods running - IaC: main.tf cleaned up — single unified rke2_worker block, worker-05 node=ms-a2-2, worker-06 removed - PCI mapping: amd-gpu now includes n5pro + ms-a2-1 + ms-a2-2 - Commit: 21bd335 - Also: ms-a2-2 upgraded PVE 9.1.1→9.2.5, kernel 7.0.14-6-pve, VFIO persistent after reboot, HA integrated ## [2026-07-24] init | ms-a2-2 Node Initialization - New Proxmox node ms-a2-2 (10.0.20.93, nodeid 1) joined cluster - Hardware: AMD Ryzen 9 9955HX (32 threads), 91GB RAM, 1.8TB NVMe - GPU: AMD Radeon 1002:13c0 (identical to ms-a2-1) — VFIO passthrough configured - PCI mapping: amd-gpu-ms-a2-2 created, added to generic amd-gpu mapping - Ceph OSD 14: created on nvme0n1 (1.8TB SSD), class ssd, weight 1.82 - Fluent-bit: installed + configured (hostname adapted to ms-a2-2) - APT sources: fixed from enterprise to no-subscription (ceph + pve) - LLM Wiki: ip-map.md, ceph.md updated; AGENTS.md updated (9 nodes) - Note: PVE 9.1.1 on ms-a2-2 (rest of cluster is 9.2.3) — upgrade pending ## [2026-07-24] cleanup | proxmox1 Permanent Removal + Ceph PG Fixes - Removed proxmox1 from Ceph CRUSH (empty host bucket, weight 0) - Also removed stale empty buckets: proxmox, px-tmp20 - Updated corosync.conf: removed proxmox1 node entry, config_version 14→15, expected_votes 9→8 - Removed stale /etc/pve/nodes/{proxmox1,proxmox,pve,px-tmp20} - Updated IaC: proxmox_api_url 10.0.20.10→10.0.20.91, template_node proxmox1→n5pro, target_nodes list updated - Updated AGENTS.md topology (9→8 nodes), LLM Wiki ip-map.md + ceph.md - Commit: 469d29d - Additional: Manual pg-upmap for stuck PGs 5.13→[1,6,7], 6.6c→[11,8,7] — all clean+remapped eliminated ## [2026-07-24] fix | Ceph Remapped PGs — Weight Imbalance + pg_num Mismatch - RCA: 11 active+clean+remapped PGs caused by CRUSH placement failure - Root cause 1: Pool 6 (hdd_disk) pg_num=120 ≠ pgp_num=112 (autoscaler mid-merge) - Root cause 2: Extreme HDD host weight imbalance (n5pro=10TB vs proxmox6=0.3TB) - Fix: Equalized pg_num→112, set nopgchange=true, reweighted osd.7 0.80→1.0 - Result: 9/11 PGs recovering (backfilling at 36MiB/s), 2 cosmetic clean+remapped remain - Created systems/ceph.md with full OSD/pool/CRUSH documentation ## [2026-07-24] update | Gitea Actions K8s Runner Bug + DinD Pitfalls - Documented Gitea 1.27.0 bug: actions_ready_job queue doesn't assign tasks to K8s-based runners - Updated systems/gitea.md with known issue, DinD pitfalls, runner management via DB - Solution doc: docs/solutions/bug-fixes/2026-07-24-gitea-actions-k8s-runner-task-assignment-bug.md - IaC committed: runner-deployment.yaml + runner-config.yaml (DinD TCP sidecar pattern) ## [2026-04-28] init | Memory System Created - Created memory directory structure at ~/.hermes/memory - Installed qmd 2.1.0 (Query Markup Documents) - Created initial index.md with entities, concepts, and sources - Created Entity pages: Infrastructure, Email-System, Health-Fitness, Team-Structure, Memory-System - Created Concept pages: Network-Architecture, Email-Organization, Monitoring-System, Proxmox-Cluster, qmd-Knowledge-Base - Initialized git repo, pushed to origin/main ## [2026-04-29] fix | Memory Process Established - Corrected stale memory-entry-001.md - Established protocol: memory update + git sync on every session ## [2026-04-30] ingest | cloud.familie-schoen.com - Full port scan, DNS, SSL, HTTP headers, API analysis - Identified: Seafile 13.0.19 on nginx, Let's Encrypt SSL ## [2026-07-24] restructure | Full Wiki Restructuring - Deleted 5 root-duplicate files (dominik-schoen, email-organization, team-structure, llm-wiki-pattern, embedding-model) - Deleted obsolete Entities/Memory-System.md, Concepts/qmd-Knowledge-Base.md - Deleted raw/articles/ and memory-entry-001/002.md (stale) - Renamed Entities/ → entities/, Concepts/ → concepts/ (lowercase) - Created new directories: systems/, reference/ - Moved Proxmox-Cluster, Monitoring-System, seafile to systems/ - Updated all existing pages with current data (April → July 2026): - Infrastructure: 2 Nodes → 9 Nodes, added Ceph/RKE2/Galera/Loki - Proxmox-Cluster: PVE 9.2.3, 9 Nodes, Fluent Bit details, PVE Tasks, Ceph Audit - Network-Architecture: cleaned up, links to reference pages - Monitoring: Prometheus/Grafana/Loki/HolmesGPT (was just "weekly network scans") - Created 6 new system pages: ceph-cluster, rke2-kubernetes, galera-maxscale, loki-fluentbit, gitea, hindsight - Created 3 reference pages: ip-map, ssh-keys, ports - Created 2 concept pages: gitops-workflow, credential-policy - Rewrote index.md with new structure + Memory Layer Architecture table - Total: 18 pages (was 21 with dupes, now 18 clean unique pages) ## [2026-09-29] update | VM302 Zombie-Recovery nach Migration - VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot - Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause - Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min - Endstand: 3/3 Synced, Primary, MaxScale alle Server Running - Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md - Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)