Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b527e4ff3c | ||
|
|
3a57db957e | ||
|
|
cc2f237ca3 | ||
|
|
531d6bfbdd | ||
|
|
a68665eca0 |
@@ -36,6 +36,7 @@
|
||||
- [[systems/frigate]] — NVR v0.18 CT151, Kameras, MQTT, HA Automation
|
||||
- [[systems/homeassistant]] — HAOS, MQTT, Automations, Notify
|
||||
- [[systems/paperless]] — Dokumentenmgmt, OIDC, OCR via noris AI
|
||||
- [[systems/sarah-hermes]] — VM107, zweites Hermes für Sarah (WebUI :8787, TG geplant)
|
||||
|
||||
## Concepts
|
||||
- [[concepts/network-architecture]] — 10.0.X.Y Schema, VLANs
|
||||
@@ -60,6 +61,11 @@
|
||||
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
|
||||
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
|
||||
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
|
||||
- [[patterns/pve-oom-frozen-guest]] — Host-OOM friert Gast als Zombie ein, RSS-Kollaps-Signatur (PAT-010)
|
||||
- [[patterns/pve-notification-webhook]] — PVE Webhook→Telegram: Endpoint-Drift, Base64-Trap, Handlebars-Escape (PAT-011)
|
||||
- [[patterns/placement-guardrails]] — 70% RAM-Deckel + Swap-Verbot als Prometheus-Rules (PAT-012)
|
||||
- [[patterns/prometheus-config-rescue]] — grep -vn Gift/Crashloop, LXC-Rootfs-Direct-Mount-Rescue (PAT-013)
|
||||
- [[patterns/ceph-dead-nvme-resurrection]] — PCI-Reset→lvchange→chown-Falle, Weight-Management bei Reintegration (PAT-014)
|
||||
- [[patterns/skill-impact]] — Skill Modification Audit Trail
|
||||
|
||||
## Memory Layer Architektur
|
||||
|
||||
@@ -1,5 +1,73 @@
|
||||
# Memory Log
|
||||
|
||||
## [2026-09-30] k8s-cp01-relief-descheduler | Workload-Migration + Descheduler-Deployment
|
||||
- cp-01 bei 93% RAM → hindsight-api + hindsight-postgres + paperless nach worker-05 migriert (cordon/delete/uncordon). cp-01 RAM 93%→43%.
|
||||
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren bis konvergiert.
|
||||
- hindsight-api Live-Pin (Sep-17-Hotfix) war nie in Git → Live-Patch nötig.
|
||||
- Descheduler v0.30.0 deployt (raw Manifests, nicht Helm — Chart hat params/args-Bug + falsche API-Version). Policy: LowNodeUtilization (30/50/20 → 70/70/60), 2m Intervall, nodeFit auf Profile-Ebene.
|
||||
- RBAC-Lektion: Descheduler braucht `list/watch` auf `namespaces` in ClusterRole.
|
||||
- ArgoCD repoURL: `ssh://git@10.0.30.200:22/...` (mit `git@` User-Prefix, wie alle anderen Apps).
|
||||
- Solution Doc + INDEX + Wiki rke2-kubernetes.md aktualisiert. Commits b955d7e→8c1ecff.
|
||||
|
||||
## [2026-09-30] ceph-monitoring-hardened | Mgr-Failover-Resistenz + SMART-Korrektur
|
||||
- Ceph-Scrape war Single-Target am aktiven mgr (10.0.20.60) → jeder mgr-Failover hätte ALLE Ceph-Alerts stumm gemacht (Standbys: 200/empty-body). Fix: Union über alle 5 mgr-Kandidaten (.50,.60,.70,.91,.92:9283) in prometheus.yml — aktiver mgr liefert, Standbys harmlos leer. TOTAL DOWN TARGETS: 0. Auch 10.0.20.70:9100 (node_exporter mit ceph-fill-collector) in Scrape-Ziele aufgenommen.
|
||||
- Zwischenfail: `aliases:`-Field in Alert-Regel ungültig (RuleNode kennt das nicht) → Prometheus Fatal. Behoben durch Entfernen + force-recreate.
|
||||
- SMART-Korrektur: Kingston SFYRDK2000G (osd.2-Träger) ist GESUND (wear 6 %, media_errors 0, spare 100 %, PoH 1469) — frühere "Austausch-Kandidatin"-These RETRAKIERT. Vorfall war Controller-/Fabric-Ebene. Neu beobachten: Temps 70/78 °C, T1-throttle 3×.
|
||||
- PAT-014 erweitert (Smart-Log-Abschnitt), ceph-cluster.md korrigiert, Commit cc2f237.
|
||||
|
||||
## [2026-09-30] ceph-osd-resurrection | NVMe-Revival osd.2 + osd.3-Weight-Drama + ubuntu-Hosts-Fix (PAT-014)
|
||||
- osd.2 (ubuntu, Kingston SFYRDK2000G) tot: NVMe-Controller state=dead, VG verschwunden, errno-5. Revival-Kette: PCI remove/rescan (Ctrl kam als nvme2 zurück!) → pvscan --cache → lvchange -ay -K (stale DM-Table) → dd-Lesetest 1,4 GB/s → chown ceph:ceph am Mapper-Device (udev-Falle) → ACTIVE, Weight 1.0, 495 GiB. Drive = Replacement-Kandidatin. *(Korrektur später am selben Tag: SMART-Audit → gesund, kein Austausch; siehe nächster Eintrag.)*
|
||||
- osd.3-Lehre: re-in mit Weight 1.0 → instant 96 % voll → backfillfull, 13 PGs blockiert. Korrekt: `crush reweight osd.3 0.05` + in. Cluster 17/17 up/in, Degraded 0,78 % fallend.
|
||||
- ubuntu-Hosts-Fix: `10.0.20.100 ubuntu` in /etc/hosts aller 8 PVE-Nodes → GUI-500 „hostname lookup failed" behoben.
|
||||
- Doc: bug-fixes/2026-09-30-ceph-osd2-nvme-resurrection-osd3-drain.md, Pattern PAT-014.
|
||||
|
||||
## [2026-09-30] watchdog-self-healed | Guardrail-Rollout begleitender Incidents (PAT-013)
|
||||
- Beim Guardrail-Rollout entdeckt: Phantom-ICMP-Targets .10/.20 lebten NOCHMAL im blackbox_icmp-Abschnitt (erste Bereinigung traf nur node_exporter-Liste) → HostUnreachableICMP-Alerts. Entfernt, 12 ICMP-Probes, alle grün.
|
||||
- Eigener Fehler: `grep -vn`-Rewrite fügte Zeilennummer-Präfixe ("1:global:") in prometheus.yml ein → Prometheus Crashloop. **Rescue-Pfad etabliert:** CT-Rootfs direkt am Host mounten (`mount /dev/rbd2 /mnt/...` — rbd2 = CT141-Disk), Fix außerhalb des pct-Kanals, Container recreation. Prometheus wieder HEALTHY, 28 Targets, DOWN=[], alle Rules ok.
|
||||
- **Lessons**: (a) NIEMALS `grep -n`-Ausgaben als Rewrite-Source verwenden; (b) LXC-Rootfs-Host-Mount = universeller Rescue-Kanal wenn pct exec zickt; (c) nach jedem Rewrite YAML-validieren BEVOR recreate.
|
||||
- Doc: bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md, Commit <SHA>.
|
||||
|
||||
## [2026-09-30] guardrails-live | Placement-Policies maschinell erzwingbar (PAT-012)
|
||||
- Zwei neue Rules in CT141 (`placement_guardrails`): PVEPlacementCeilingBreached (>70% RAM, 15m) + HVResidentSwapNonzero (>100MiB, 10m). Alle Rules health=ok.
|
||||
- Baseline: proxmox3 80,6% / p4 79,3% / p5 75,5% / p6 72,6% ÜBER Deckel → Warnbursts erwartet (Rebalancing-Backlog Richtung ms-a2-1/-2 mit 74%/Headroom).
|
||||
- Doc: architecture/2026-09-30-placement-guardrails-ram-ceiling.md, Commit 66f2172.
|
||||
|
||||
## [2026-09-30] webhook-fixed | PVE→Telegram Notification-Pipeline repariert (PAT-011)
|
||||
- Ursachenkette (dreifach gestapelt): Endpoint-Drift .99→.141 (Bridge wohnt in CT141), fehlender `body`-Attr (Leere Posts → 400), unescapte Handlebars-Interpolation (Apostrophe/Multiline → invalides JSON).
|
||||
- Fix: URL korrigiert, Body via pvesh (BASE64-Pflicht!) mit `{{escape title}}`/`{{escape message}}` (Space-Syntax, NICHT Colon). Offizieller Test grün, Bridge loggt POST /pve 200.
|
||||
- Diagnose-Technik: Mini-Sniffer (temp URL-Redirect, Bytes kapern, URL restaurieren) enthüllte exakten Wire-Body.
|
||||
- Doc: bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md, Commit 38622ff. Residue ge cleaned (HV+/tmp+lokales).
|
||||
|
||||
## [2026-09-30] legacy-alert-cleanup | Alerts bereinigt, mysqld_exporter VM300 nachgezogen
|
||||
- proxmox3 /boot/efi 100%: 17 alte Kernel-Pakete gepurged (-17/-19/-12 behalten) → 30%.
|
||||
- Phantom-Targets .10/.20 entfernt, ICMP→TCP-Probe für Offsite-PBS (ICMP upstream gefiltert) → 30 Targets.
|
||||
- VM300 mysqld_exporter 0.15.1 nachdeployt (war aspirational Target): GitHub-Download via qm guest exec+b64, exporter-User (vorgeschädigte exporter@localhost-Shadow!) PW-Align, UFW 9104←10.0.30.141. Alle 3 Galera-Exporte UP.
|
||||
- Cleanup: alle /tmp-Skripts (HV+CT141+VM300), lokale Scratch-Dirs entfernt.
|
||||
|
||||
## [2026-09-30] remediation-complete | Vorfall-Nacharbeiten: GPU restored, PSI/OOM-Alerts live, p6 entlastet
|
||||
- GPU (VM102/ms-a2-2): hostpci0 ohne x-vga restauriert → renderD128 lebt. **x-vga=1 bricht AMD-Passthrough** (SeaBIOS Shadow-ROM → VBIOS-Zugriff tot, amdgpu -22). Doc-Update in 2026-07-21-amd-gpu-passthrough-rombar.md.
|
||||
- Monitoring (CT141): Regelgruppe `pressure_alerts` (5 Regeln: PSI mem waiting/stalled, oom_kill, Swap-Churn, MajFault-Storm) live; node_exporter auf ms-a2-1/-2 nachinstalliert (Blindspots!), 32 Targets. **LXC-Bindmount-Inode-Trap:** sed-i/In-Place-Rewrites unsichtbar für Docker bis Force-Recreate → neues Doc.
|
||||
- proxmox6: VM301 → proxmox7 (HA relocate, na-vm301 live verifiziert). Swap 6,3G→0B, RAM 8,4/15G. Nur noch CT151 auf p6. Altlast-Alerts sichtbar geworden (NodeDown .10/.20 Phantoms, proxmox3 /boot/efi 100%, ICMP 213.95.54.60) — Cleanup offen.
|
||||
- Docs: bug-fixes/2026-09-30-prometheus-lxc-bindmount-inode-trap.md (neu), INDEX.md aktualisiert.
|
||||
|
||||
## [2026-09-30] incident-fix | worker-05 Freeze: proxmox6 Doppel-OOM → Zombie-VM (PAT-010)
|
||||
- Symptom: KubeDaemonSetRolloutStuck (Traefik DS misscheduled=1), worker-05 NotReady 19h
|
||||
- Root Cause: proxmox6 RAM-Oversubscription (~47.8 GB alloc / 16 GB, 7.3/8 GB Swap) → OOM-Killer tötete kvm (VM102) 2× (29.09. 18:43 + 20:44); 2. Revival = Zombie (QEMU running, Gast inert, RSS 239MB/12GB)
|
||||
- Diagnose-Signatur: `qm status --verbose` RSS-Kollaps + tote Guest-Agent + statischer Tap-TX
|
||||
- Fix: Laya CT152 → proxmox3 (Offline-Move 2s, shared RBD) → Druck raus; qm stop/start VM102 → HA replatzierte auf ms-a2-2 (60 GB frei, GPU-fähig!); uncordon → Node Ready, Traefik-Orphan self-reconciled, Alert cleared
|
||||
- Drift-Fund: frühere node-affinity `na-vm102` (NICHT ms-a2-2) existiert live nicht mehr
|
||||
- Offen: proxmox6 strukturell eng (28 GB alloc); PSI/OOM-Alerts pro PVE-Node fehlen komplett; GPU(renderD128)-Verifikation auf ms-a2-2 nach Return
|
||||
- Docs: docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md, patterns/pve-oom-frozen-guest.md (PAT-010)
|
||||
|
||||
## [2026-09-30] deployment | Sarah-Hermes VM107 — zweite Hermes-Instanz live
|
||||
- VM107 (n5pro, 10.0.30.66) via Tofu epic-8 deployed; Docker+UFW via Ansible (epic-7-Stil)
|
||||
- hermes-webui Single-Container :8787, LAN-only, Password-Auth; Secrets in 1P (sarah-hermes-webui, sarah-hermes-noris-key)
|
||||
- E2E verifiziert: Login 200, Agent-Antwort via vllm/release/glm-5-2 @ ai.noris.de
|
||||
- PITFALL: hermes-agent pyproject pinnt Core-Deps hinter python_version>='3.14'-Markers → auf Python 3.12 installiert `pip install -e .` NULL Deps ("AIAgent not available"). Fix: manuelle Pin-Installation + Missing-Import-Loop bis `import run_agent`
|
||||
- Backup: täglich 23:00 Job backup-2e8a34e3-66cb (all=1) deckt VM107; manueller Verify TASK OK 48s
|
||||
- Wiki: systems/sarah-hermes.md, ip-map.md erweitert; iac-homelab commits 0d5f017/1ce4816
|
||||
- OFFEN: Telegram-Bot blockiert auf BotFather-Token (Sarah/Dominik)
|
||||
|
||||
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
|
||||
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
|
||||
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# PAT-014 — Dead-NVMe-Resurrection für Ceph-OSD
|
||||
|
||||
## Trigger
|
||||
Ceph-OSD auf externem/non-Corosync-Host startet nicht: KernelDevice errno-5 I/O-Errors beim
|
||||
Label-Lesen, Symlink `/var/lib/ceph/osd/ceph-*/block` ins Leere, VG „verschwindet".
|
||||
|
||||
## Root Cause
|
||||
NVMe-Controller im PCIe-Fabric gestorben (state=dead, Namespace 0B) — NICHT das Medium selbst.
|
||||
Nach PCI-reset kehrt der Controller oft unter NEUER Nummer zurück (nvme0 → nvme2!), alte
|
||||
Device-Mapper-Tabelle bleibt stale.
|
||||
|
||||
## Resurrection-Sequence (verifizierte Reihenfolge)
|
||||
1. **Korrekte PCI-Addr finden**: über sysfs-Pfad des Namespaces (`/sys/class/block/nvmeXnY/device`),
|
||||
NICHT raten — erster Versuch traf den falschen (gesunden!) Controller.
|
||||
2. `echo 1 > /sys/bus/pci/devices/<ADDR>/remove && echo 1 > /sys/bus/pci/rescan`
|
||||
→ Controller kommt als neue Instanz zurück, Namespaces wieder da.
|
||||
3. `pvscan --cache` → VG/LV wieder sichtbar.
|
||||
4. **Stale DM-Table**: `lvchange -an <vg>/<lv> && lvchange -ay -K <vg>/<lv>`
|
||||
5. **Lesetest**: `dd if=/dev/<vg>/<lv> bs=4M count=8 of=/dev/null` — bei weiteren errno-5:
|
||||
NAND/Media defekt → OSD out lassen, Drive tauschen.
|
||||
6. **Permissions-Falle nach lvchange**: udev setzt Owner root →
|
||||
`chown ceph:ceph /dev/mapper/<dm-name>; chmod 660` (sonst `bdev open: (13) Permission denied`).
|
||||
7. `systemctl reset-failed ceph-osd@<id> && systemctl start ceph-osd@<id>`
|
||||
8. `ceph osd in osd.<id>` + Weight restaurieren.
|
||||
|
||||
## Begleitregeln
|
||||
- Klein/niedergewichtetes OSD niemals mit Weight 1.0 reintegrieren, solange Pools an
|
||||
Full-Ratios kratzen → sonst `backfillfull`/`backfill_toofull`-Blockade (Fall osd.3:
|
||||
re-in@1.0 → 96 % voll instant → 13 PGs blocked; Fix: `crush reweight osd.3 0.05`).
|
||||
- Externe Non-Corosync-Hosts in `/etc/hosts` ALLER PVE-Nodes pflegen (GUI-500
|
||||
„hostname lookup failed"), oder echter DNS-Record.
|
||||
|
||||
## Verified
|
||||
2026-09-30: osd.2 revived (Kingston SFYRDK2000G, PCI 03:00.0, → nvme2), 495 GiB, Weight 1.0;
|
||||
Cluster 17/17 up/in; Degraded 0,78 % fallend. osd.3 stabilized @ weight 0.05.
|
||||
|
||||
## Smart-Log-Abgleich 2026-09-30 (korrigiert frühere Alters-These)
|
||||
Kingston SFYRDK2000G (nvme0n2, trägt osd.2): percentage_used **6 %**, media_errors **0**,
|
||||
available_spare 100 %, unsafe_shutdowns 2, power_on_hours 1469, power_cycles 3.
|
||||
Geschwisterplatte nvme1n1 identisches Profil (6 % wear, 0 media errors).
|
||||
⇒ Drive ist MEDIZINISCH GESUND — der Vorfall war rein Controller-(PCI)-Ebene, kein
|
||||
Media-Verschleiß. **Kein Austausch nötig.** Einziges Beobachtungsfeld: Temperatur 70 °C /
|
||||
Sensor2 78 °C (thermisches Throttling T1 3× aktiviert) — Kühlung prüfen wäre sinnvoll,
|
||||
aber keine Akutgefahr.
|
||||
@@ -0,0 +1,23 @@
|
||||
# PAT-012 — Placement Guardrails (RAM-Deckel + Swap-Verbot)
|
||||
|
||||
**Trigger**: Neue Gastplatzierung oder Kapazitätsfrage auf einem PVE-Node.
|
||||
|
||||
**Policy**:
|
||||
- Committed RAM pro Hypervisor ≤ 70 % physikalisch. Über Deckel = "voll",
|
||||
Guests redistribuieren BEVOR neue Last kommt.
|
||||
- Resident Swap > 100 MiB für >10 min = Politikverstoß → Placements prüfen,
|
||||
nie als Normalzustand akzeptieren.
|
||||
|
||||
**Enforcement (live)**: Gruppe `placement_guardrails` in CT141
|
||||
(`/opt/monitoring/prometheus/rules/alerting_rules.yml`):
|
||||
- `PVEPlacementCeilingBreached` (ratio > 0.70, for 15m, warning)
|
||||
- `HVResidentSwapNonzero` (SwapTotal-Free > 100MiB, for 10m, warning)
|
||||
|
||||
**Vor jeder Platzierung**: Ratio-Query ziehen
|
||||
(`pve_memory_usage_bytes{id=~"node/.*"} / pve_memory_size_bytes{id=~"node/.*"}`)
|
||||
— >70 %-Nodes meiden, Headroom liegt primär auf ms-a2-1/ms-a2-2.
|
||||
|
||||
Baseline-Rollout 30.09.: proxmox3 80,6 / p4 79,3 / p5 75,5 / p6 72,6 %
|
||||
über Deckel (Warnburst erwartet = Rebalancing-Backlog, kein Emergency).
|
||||
|
||||
Ref: docs/solutions/architecture/2026-09-30-placement-guardrails-ram-ceiling.md
|
||||
@@ -0,0 +1,28 @@
|
||||
# PAT-013 — Prometheus Config Rescue (Crashloop + LXC Rootfs Channel)
|
||||
|
||||
**Trigger**: Prometheus crashloop nach Config-Rewrite, oder pct exec unbrauchbar.
|
||||
|
||||
**Iron Rules**:
|
||||
1. NIEMALS `grep -n`-Ausgabe als Rewrite-Source — Nummernpräfixe verseuchen die Datei.
|
||||
Nutze `grep -v` (ohne -n) oder Python/awk.
|
||||
2. YAML validieren BEVOR Container recreate (sonst Crashloop = totale Blindheit).
|
||||
3. Beim Löschen von Targets/IPs ALLE Locations sweeppen — Dubletten leben gern in
|
||||
mehreren Jobs derselben Datei (node_exporter-Liste ≠ blackbox-Liste).
|
||||
|
||||
**Rescue-Kanal (universell für LXC auf RBD)**:
|
||||
```bash
|
||||
# Auf dem Hyper visor, der den CT hostet:
|
||||
lsblk # rbdX finden (size matchen)
|
||||
mount /dev/rbdX /mnt/rescue
|
||||
# Dateien direkt editieren unter /mnt/rescue/opt/...
|
||||
umount /mnt/refuge
|
||||
# im CT: docker compose up -d --force-recreate <svc>
|
||||
```
|
||||
|
||||
**Symptom-Signature Crashloop**: `docker logs prometheus` → "yaml: line N: mapping
|
||||
values are not allowed in this context" = klassisches NN:-Präfix-Gift.
|
||||
|
||||
**Stale-Series-Note**: gelöschte Targets erscheinen sekundenweise weiter in Queries —
|
||||
erst nach Scrape-Zyklus-Gap re-checken, dann Erfolg erklären.
|
||||
|
||||
Ref: docs/solutions/bug-fixes/2026-09-30-prometheus-yaml-prefix-crashloop.md
|
||||
@@ -0,0 +1,22 @@
|
||||
# PAT-011 — PVE Webhook Notifications Pipeline Repair
|
||||
|
||||
**Trigger**: PVE notifications (esp. vzdump) reach Telegram not / test returns 500.
|
||||
|
||||
**Signature diagnosis ladder**:
|
||||
1. `pvesh create /cluster/notifications/targets/<name>/test` — error taxonomy:
|
||||
- `Connection refused` → transport/down (check URL target alive)
|
||||
- `failed to render webhook body` → template broken (syntax/base64)
|
||||
- `http status: 400` → bridge rejected payload (escaping!)
|
||||
2. Listener-Sweep über VLAN: `for ip in .xx…; do /dev/tcp/$ip/port probe; done`
|
||||
3. Byte-Level-Truth via temp mini-sniffer (redirect URL, capture, restore!).
|
||||
|
||||
**Hard rules**:
|
||||
- Mutations an `/etc/pve/notifications.cfg` NUR via `pvesh set` (hand-edits poison
|
||||
global deserialization!). Vorher Snapshot.
|
||||
- `body`/header-values/secrets = **base64 blobs**: `printf '%s' tpl | base64 -w0`.
|
||||
- Handlebars helpers: `{{escape title}}` (Space!), NICHT `{{escape:title}}`.
|
||||
- Immer `{{escape title}}`/`{{escape message}}` verwenden — rohe Interpolation bricht
|
||||
bei Apostrophen/Multiline (Backup-Reports!).
|
||||
- Sniffer-Redirect URL IMMER restaurieren.
|
||||
|
||||
Reference: `docs/solutions/bug-fixes/2026-09-30-pve-webhook-notifications-drift-base64-escape.md`
|
||||
@@ -0,0 +1,49 @@
|
||||
---
|
||||
pattern_id: PAT-010
|
||||
title: "PVE host OOM-kill freezes guest VM as zombie (RSS collapse signature)"
|
||||
category: infrastructure
|
||||
severity: high
|
||||
status: active
|
||||
first_observed: 2026-09
|
||||
last_updated: 2026-09-30
|
||||
related_systems: [proxmox-cluster, rke2-kubernetes]
|
||||
related_solution_docs: [docs/solutions/bug-fixes/2026-09-30-proxmox6-oom-frozen-vm102-worker05.md]
|
||||
related_skills: [rke2-cluster-administration, systematic-debugging]
|
||||
---
|
||||
|
||||
# PAT-010: PVE host OOM-kill freezes guest VM as zombie
|
||||
|
||||
## Symptom
|
||||
- K8s node NotReady, kubelet heartbeat stops abruptly
|
||||
- PVE shows VM `running`, but: no ping, no SSH, qemu-guest-agent dead,
|
||||
tap-interface TX counters static
|
||||
- **Signature:** `qm status <vmid> --verbose` → kvm RSS collapses to a
|
||||
tiny fraction (<5%) of assigned RAM — guest kernel no longer touches
|
||||
its memory
|
||||
- Downstream alerts (DS misscheduled, workload CrashLoops) fire, but
|
||||
NOTHING alarms on the actual OOM
|
||||
|
||||
## Root Cause
|
||||
Host RAM oversubscription (allocations >> physical RAM). OOM-killer picks
|
||||
the largest anon-RSS process = biggest kvm. After TWO consecutive kills of
|
||||
the same guest (HA auto-restarts in between), the revived QEMU comes up
|
||||
but the guest kernel stays inert → zombie VM.
|
||||
|
||||
## Mitigation
|
||||
1. Relieve host pressure FIRST (move movable tenants away) — otherwise
|
||||
the unfreeze re-boots the guest into the same thrash.
|
||||
2. Hard cycle the VM: `qm stop` + `qm start` (respect HA guards).
|
||||
3. Re-query `ha-manager status` afterwards — HA may relocate the VM.
|
||||
4. Uncordon the K8s node; orphan DS pods self-reconcile.
|
||||
|
||||
## Prevention
|
||||
- Allocation ceiling per PVE host (≤ ~70% of RAM) enforced in placement/
|
||||
IaC; swap is a buffer, not capacity.
|
||||
- PSI/OOM alerting per PVE node in Prometheus — OOM kills are currently
|
||||
invisible to alerting (noticed only via downstream K8s symptoms).
|
||||
|
||||
## Evidence
|
||||
- 2026-09-30: proxmox6 (16 GB, ~47.8 GB allocated): OOM-killed
|
||||
rke2-worker-05 kvm twice (18:43:38, 20:44:12 UTC on 29.09.), second
|
||||
revival = zombie (RSS 239 MB / 12 GB). Fixed via Laya-CT migration +
|
||||
hard recycle; HA relocated VM to ms-a2-2. Full RCA in related doc.
|
||||
+2
-1
@@ -3,7 +3,7 @@ title: IP-Map (Quick Reference)
|
||||
category: reference
|
||||
tags: [ip, network, reference, quick-lookup]
|
||||
created: "2026-07-24"
|
||||
modified: "2026-09-26"
|
||||
modified: "2026-09-30"
|
||||
---
|
||||
|
||||
# IP-Map
|
||||
@@ -47,6 +47,7 @@ modified: "2026-09-26"
|
||||
## Infrastructure VMs (10.0.30.x)
|
||||
| IP | Host | Service |
|
||||
|----|------|---------|
|
||||
| 10.0.30.66 | sarah-hermes (VM107, n5pro) | Sarahs Hermes: WebUI :8787 (LAN-only, pw) + TG-Bot (geplant) |
|
||||
| 10.0.30.99 | CT111 | Immich (n5pro), Migration zu K8s geplant |
|
||||
| 10.0.30.100 | ubuntu | Physischer Ubuntu-Node (Ceph OSDs, ZFS pool01_n2_redundant) |
|
||||
| 10.0.30.105 | — | ~~Gitea~~ CT108 GESTOPPT 2026-09-17 (Archive: /home/debian/git-archive/ct108-final/) |
|
||||
|
||||
+13
-1
@@ -23,7 +23,7 @@ modified: "2026-09-26"
|
||||
|-----|-------|------|------|----------|-------|
|
||||
| 1 | hdd | 3.7 TiB | n5pro | 1.0 | |
|
||||
| 2 | ssd | 1.8 TiB | ubuntu | 1.0 | Moved to ubuntu host |
|
||||
| 3 | ssd | 233 GB | proxmox4 | 1.0 | |
|
||||
| 3 | ssd | 233 GB | proxmox4 | 0.05 | 2026-09-30: reweighted 0.05 nach Full-Drama (war 1.0) — Plate 238G, sonst backfillfull |
|
||||
| 4 | ssd | 227 GB | proxmox3 | 0.30 | Small, reweighted down |
|
||||
| 5 | ssd | 150 GB | proxmox5 | 0.30 | Small, reweighted down |
|
||||
| 6 | hdd | 3.6 TiB | n5pro | 1.0 | |
|
||||
@@ -73,6 +73,18 @@ Fix: `ceph config set mon mon_allow_insecure_global_id_reclaim false` (if not al
|
||||
osd.4 (227GB, proxmox3) and osd.5 (150GB, proxmox5) reweighted to 0.30 — too small for meaningful capacity.
|
||||
Consider removing from CRUSH or replacing with larger drives.
|
||||
|
||||
### NVMe-Controller-Death auf ubuntu + Recovery (2026-09-30, PAT-014)
|
||||
Kingston SFYRDK2000G (PCI 03:00.0) starb (state=dead, VG verschwand) → osd.2 down/out.
|
||||
Revived via PCI remove/rescan + lvchange -ay -K + chown-Falle am mapper-device.
|
||||
Details: patterns/ceph-dead-nvme-resurrection (PAT-014). **Update 2026-09-30 (Abend): SMART-
|
||||
Audit spricht FREI — percentage_used 6 %, media_errors 0, spare 100 %, PoH 1469. Vorfall war
|
||||
rein Controller-Ebene, kein Media-Verschleiß, KEIN Austausch nötig. Beobachten: Temps 70/78 °C,
|
||||
thermal throttle T1 3×.**
|
||||
|
||||
### ubuntu-Host in /etc/hosts aller PVE-Nodes (2026-09-30)
|
||||
Ohne DNS-Record wirft die PVE-GUI `hostname lookup 'ubuntu' failed (500)`.
|
||||
Fix: hosts-Eintrag `10.0.20.100 ubuntu` fleetweit auf allen 8 Nodes.
|
||||
|
||||
### worker-04 (VM 139) NotReady in K8s
|
||||
Node offline — not a Ceph issue but affects Ceph CSI attachments.
|
||||
|
||||
|
||||
@@ -3,7 +3,7 @@ title: Proxmox VE Cluster
|
||||
category: systems
|
||||
tags: [proxmox, virtualization, lxc, qemu, pve]
|
||||
created: "2026-04-28"
|
||||
modified: "2026-09-29"
|
||||
modified: "2026-09-30"
|
||||
---
|
||||
|
||||
# Proxmox VE Cluster
|
||||
@@ -50,13 +50,24 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
||||
- Benötigte modprobe.d Config:
|
||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||
- Worker-05 (VM 102) lief auf ms-a2-2 mit funktionierendem GPU-Passthrough. Seit 2026-09-29 auf proxmox6 (Anti-Collocation mit Worker-01). Prüfen ob GPU-Passthrough auf proxmox6 funktioniert!
|
||||
- Worker-05 (VM 102): GPU-Passthrough seit 2026-09-30 WIEDER AKTIV auf ms-a2-2 —
|
||||
`hostpci0: 0000:01:00.0,pcie=1,rombar=1` (OHNE x-vga!) → renderD128 verifiziert.
|
||||
**Kritische Lehre:** `x-vga=1` bricht moderne AMD-Karten (SeaBIOS Shadow-ROM zerstört
|
||||
VBIOS-Zugriff, amdgpu error -22 "Unable to locate a BIOS ROM"). Für Headless-
|
||||
Render-Nodes NIEMALS x-vga kombinieren. Frühere node-affinity `na-vm102` existiert
|
||||
live NICHT (affinity.cfg verifiziert 30.09.) — nur resource-affinity vs 128/139.
|
||||
|
||||
## Bekannte Probleme
|
||||
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||
- osd.5 reweight 0.30 (kleine SSD, 150GB) — entlasten oder austauschen
|
||||
- osd.4 reweight 0.30 (kleine SSD, 227GB auf proxmox3) — gleiche Situation
|
||||
- worker-04 (VM 139) NotReady im K8s Cluster — Node offline
|
||||
- **proxmox6 RAM-Oversubscription (AKUT ENTSCHÄRFT 30.09.):** VM301 (Galera db2,
|
||||
8G) am 30.09. via `ha-manager relocate vm:301 proxmox7` migriert (na-vm301 =
|
||||
5/6/7 verifiziert; Anti-Collocs 300⊥301, 301⊥302 gewahrt). Danach: 8,4/15G RAM,
|
||||
Swap 6,3G→2,8G, per swapoff/on geleert → 0B. Verbleibt auf p6: nur CT151
|
||||
Frigate (8G) — innerhalb Ceiling. TODO bleibt: Placement-Ceiling (~70%) als
|
||||
Guardrail formalisieren. PSI/OOM-Alerts: LIVE in CT141 (Regelgruppe
|
||||
`pressure_alerts`, 5 Regeln; node_exporter nachinstalliert auf ms-a2-1/-2).
|
||||
|
||||
## HA Rules (PVE 9.2 Rules System)
|
||||
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||
@@ -112,10 +123,10 @@ Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
||||
| RKE2 CP-03 | 126 | proxmox4 |
|
||||
| RKE2 Worker-01 | 128 | proxmox5 |
|
||||
| RKE2 Worker-04 | 139 | n5pro |
|
||||
| RKE2 Worker-05 | 102 | proxmox6 |
|
||||
| RKE2 Worker-05 | 102 | ms-a2-2 (seit 30.09.; vorher proxmox6, davor ms-a2-2) |
|
||||
| Hermes-Agent-01 | 230 | n5pro |
|
||||
| Galera db1 | 300 | n5pro |
|
||||
| Galera db2 | 301 | proxmox6 |
|
||||
| Galera db2 | 301 | **proxmox7** (seit 30.09. relocate; vorher proxmox6) |
|
||||
| Galera db3 | 302 | ms-a2-2 |
|
||||
| MaxScale-01 | 310 | proxmox7 |
|
||||
| MaxScale-02 | 311 | ms-a2-1 |
|
||||
|
||||
@@ -22,7 +22,7 @@ modified: "2026-09-17"
|
||||
| cp-03 | 10.0.30.53 | Control Plane |
|
||||
| worker-01 | 10.0.30.63 | Worker |
|
||||
| worker-04 | 10.0.30.64 | Worker |
|
||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2; 30.09. OOM-Zombie-Freeze behoben, siehe PAT-010) |
|
||||
|
||||
## Storage
|
||||
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
||||
@@ -66,6 +66,19 @@ modified: "2026-09-17"
|
||||
- IaC: unified `amd-gpu` PCI mapping (n5pro + ms-a2-1 + ms-a2-2) + dynamic hostpci (one worker block, gpu flag)
|
||||
- ms-a2-1 has kernel BUG with 1002:13c0 (renderD128 missing) — worker-05 moved to ms-a2-2 (identical hardware)
|
||||
|
||||
## Capacity Management
|
||||
- **Descheduler** v0.30.0 deployed als ArgoCD-App (`clusters/main/descheduler/manifests.yaml`)
|
||||
- Plugin: `LowNodeUtilization` (thresholds 30/50/20, targetThresholds 70/70/60)
|
||||
- Interval: 2m (`--descheduling-interval=2m`)
|
||||
- `nodeFit: true` auf Profile-Ebene (nicht Plugin-Ebene — v0.30 Schema!)
|
||||
- RBAC: benötigt `list/watch` auf `namespaces` (fehlt in Default-Helm-Chart)
|
||||
- Raw Manifests statt Helm Chart (Chart 0.30 hat params/args-Bug, Chart 0.35 hat falsche API-Version)
|
||||
- Siehe Solution Doc `2026-09-30-k8s-cp01-relief-descheduler.md`
|
||||
- **CP-01 Relief (2026-09-30):** hindsight-api + hindsight-postgres + paperless von cp-01 → worker-05 migriert
|
||||
- cp-01 RAM: 93% → 43%; worker-05 bei ~28%
|
||||
- ArgoCD selfHeal belebte alte ReplicaSets mit nodeSelector wieder → manuell auf 0 skalieren
|
||||
- Live-only nodeSelector (hindsight-api, Sep-17-Hotfix) war nie in Git → Live-Patch nötig
|
||||
|
||||
## Known Pitfalls
|
||||
- `enableServiceLinks: false` bei Apps deren Service-Name mit Env-Vars kollidiert (z.B. Paperless `PAPERLESS_PORT`)
|
||||
- ArgoCD `--force` kann nicht mit ServerSideApply kombiniert werden
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
---
|
||||
title: Sarah-Hermes (VM107)
|
||||
category: systems
|
||||
tags: [hermes, webui, vm, sarah, telegram, backup]
|
||||
created: "2026-09-30"
|
||||
modified: "2026-09-30"
|
||||
related: [systems/rke2-kubernetes, reference/ip-map, systems/noris-ai]
|
||||
---
|
||||
|
||||
# Sarah-Hermes (VM107)
|
||||
|
||||
Zweite, vollständig isolierte Hermes-Instanz für Sarah (Allround-Assistentin:
|
||||
Erinnerungen, Planung, Smalltalk). Getrennte Memories/Sessions/Skills/Keys von
|
||||
den 5 Owner-Profilen.
|
||||
|
||||
## Eckdaten
|
||||
|
||||
| Attribut | Wert |
|
||||
|----------|------|
|
||||
| VM-ID | 107 (auto-allokiert) |
|
||||
| Node | n5pro (Template 9000 debian-12-cloudinit) |
|
||||
| IP | 10.0.30.66/24 (static via DHCP reservation) |
|
||||
| Specs | 2 vCPU / 4 GB RAM / 32 GB Disk (vm_disks/RBD) |
|
||||
| Access | SSH `debian@10.0.30.66` mit `~/.ssh/id_ed25519_cloudinit` |
|
||||
| Tofu | `iac-homelab/epic-8-sarah-hermes/tofu/` |
|
||||
| Ansible | `iac-homelab/epic-8-sarah-hermes/ansible/` |
|
||||
| Plan | `iac-homelab/docs/plans/2026-09-30-sarah-hermes-vm.md` |
|
||||
|
||||
## Stack
|
||||
|
||||
- **hermes-webui** (ghcr.io/nesquena/hermes-webui:latest), Single-Container,
|
||||
Port 8787 (0.0.0.0 gebunden, UFW erlaubt nur LAN), Password-Auth.
|
||||
Compose: `/home/debian/hermes-webui/docker-compose.yml` auf der VM.
|
||||
- **Agent-Runtime:** nousresearch/hermes-agent geklont nach
|
||||
`~/.hermes/hermes-agent` auf der VM; installiert in `/app/venv` im Container
|
||||
(editable). **Wichtig:** Core-Deps im pyproject sind hinter
|
||||
`python_version >= '3.14'`-Markern gepinnt → auf Python 3.12 installiert
|
||||
`-e .` NULL Deps. Manual-Dep-Bootstrapping nötig (siehe Pitfalls).
|
||||
- **LLM:** noris-Provider (`https://ai.noris.de/v1`, Default-Modell
|
||||
`vllm/release/glm-5-2`), Key via `HERMES_CUSTOM_NORIS_API_KEY` aus
|
||||
`.env` (Compose mapped explizit in den Container).
|
||||
- **Telegram:** geplant (blockiert auf BotFather-Token von Sarah/Dominik).
|
||||
Bei Aktivierung: `gateway.telegram_enabled: true` + Token in `.env`;
|
||||
Webhook/API-Server-Ports bleiben disabled (Konfliktvermeidung).
|
||||
|
||||
## Secrets (1Password, Vault: Hermes)
|
||||
|
||||
- `sarah-hermes-webui` → HERMES_WEBUI_PASSWORD
|
||||
- `sarah-hermes-noris-key` → HERMES_CUSTOM_NORIS_API_KEY
|
||||
|
||||
## Backup
|
||||
|
||||
- Daily vzdump-Job `backup-2e8a34e3-66cb` (23:00, `all=1`, exclude 301,302,310,311,137,147,501)
|
||||
→ **deckt VM107 ab** (Storage `noris_v4` = PBS Datastore `noris` @ 10.0.30.119).
|
||||
- Manueller Verify-Lauf am 30.09.: TASK OK in 48s (inkrementell, 88% reuse).
|
||||
- Offsite: folgt dem regulären `push-offsite` Sync-Job (Pull↔Push-Korrektur
|
||||
vom 19.09.).
|
||||
|
||||
## Known Issues / Pitfalls
|
||||
|
||||
1. **Python-Version-Mismatch:** hermes-agent pyproject pins Core-Deps an
|
||||
`python_version >= '3.14'`; WebUI-Container läuft auf 3.12 → `-e .`
|
||||
installiert keine Deps. Fix: manuell `pip install` der gepinschten Pakete
|
||||
+ iterativer Missing-Import-Loop bis `import run_agent` klappt.
|
||||
2. **hermes update Ownership-Konflikt:** `hermes update`-Completion beschwert
|
||||
sich über uid 0 vs. uid 1000 auf `/app/venv/bin/hermes-acp`. Kosmetisch,
|
||||
Betriebsbetrieb unbeeinträchtigt. Fix-Idee: `chown` im Entry-Point.
|
||||
3. **telegram-bridge Notify-Target defekt** (seit 19.09., Connection refused):
|
||||
betrifft vzdump-Notifications clusterweit, nicht nur VM107. Separater Fix.
|
||||
|
||||
## Verification History
|
||||
|
||||
- 2026-09-30: Deploy + E2E-Test (Login 200, Agent-Antwort "HALLO" via glm-5-2).
|
||||
- 2026-09-30: Manueller vzdump → TASK OK (48s, inkrementell).
|
||||
Reference in New Issue
Block a user