Auto-sync: 2026-09-29
This commit is contained in:
@@ -1,5 +1,14 @@
|
|||||||
# Memory Log
|
# Memory Log
|
||||||
|
|
||||||
|
## [2026-09-29] bug-fix | KubeAPIErrorBudgetBurn — Ceph-CSI resize loop from missing controller-expand-secret
|
||||||
|
- Root Cause: `ceph-flash` + `ceph-hdd-replica` StorageClasses created manually WITHOUT `controller-expand-secret-name/namespace` params
|
||||||
|
- CNPG PVC resize (50→100Gi, 10→20Gi) triggered infinite CSI resizer retry loop ("provided secret is empty") → API server write pressure → etcd DeadlineExceeded → Handler timeout 5xx
|
||||||
|
- Cordoning cp-03 amplified: CNPG switchover → operator reconcile storm (~60/min)
|
||||||
|
- Fix: (1) Delete+recreate SCs with all secret refs, (2) Patch 6 PVs with controllerExpandSecretRef, (3) Restart CSI resizer, (4) Commit SCs to Git `clusters/main/storage/`, (5) Move Gitea off unstable worker-05
|
||||||
|
- Worker-05 cordoned (repeated reboots, likely hypervisor issue on ms-a2-2)
|
||||||
|
- Architectural risk documented: etcd on Ceph RBD (~20ms WAL fsync all CP nodes)
|
||||||
|
- Solution doc: docs/solutions/bug-fixes/2026-09-29-kubeapi-error-budget-burn-ceph-csi-resize-loop.md
|
||||||
|
|
||||||
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
|
## [2026-09-27] bug-fix | Gitea CSI RBAC Fix + ArgoCD Verknüpfungs-Audit
|
||||||
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
|
- Root Cause: Ceph CSI RBD `csi-attacher` fehlte `storage.k8s.io/csinodes` Berechtigung → VolumeAttachments pending → Gitea Pod 4+ Tage Init-crash
|
||||||
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
|
- Fix: ClusterRole `ceph-rbd-external-attacher-runner` patched, provisioner Pods neu gestartet → alle 24 VolumeAttachments `true`
|
||||||
@@ -213,3 +222,11 @@
|
|||||||
- Created 2 concept pages: gitops-workflow, credential-policy
|
- Created 2 concept pages: gitops-workflow, credential-policy
|
||||||
- Rewrote index.md with new structure + Memory Layer Architecture table
|
- Rewrote index.md with new structure + Memory Layer Architecture table
|
||||||
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
|
- Total: 18 pages (was 21 with dupes, now 18 clean unique pages)
|
||||||
|
|
||||||
|
## [2026-09-29] update | VM302 Zombie-Recovery nach Migration
|
||||||
|
- VM302 (Galera db3) reagierte nach Migration auf ms-a2-2 nicht: QEMU "running", aber SSH/MariaDB/QGA tot
|
||||||
|
- Root Cause: Post-Migration-Zombie; -incoming/-S in QEMU-Cmdline ist Artefakt, kein Beweis für Pause
|
||||||
|
- Fix: qm stop/start trotz HA-Guard, SST-Rejoin ~2-3min
|
||||||
|
- Endstand: 3/3 Synced, Primary, MaxScale alle Server Running
|
||||||
|
- Solution Doc: docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md
|
||||||
|
- Zusätzlich: Home Assistant Core Restart via REST API erfolgreich (Version 2026.9.3, RUNNING)
|
||||||
|
|||||||
@@ -35,6 +35,11 @@ modified: "2026-07-24"
|
|||||||
- VM300/301/302 (Galera): Fluent Bit aktiv
|
- VM300/301/302 (Galera): Fluent Bit aktiv
|
||||||
- VM310/311 (MaxScale): Fluent Bit aktiv
|
- VM310/311 (MaxScale): Fluent Bit aktiv
|
||||||
|
|
||||||
|
## Bekannte Issues
|
||||||
|
- **2026-09-29:** VM302 Zombie-State nach Migration (QEMU running, Services tot, QGA down). Behoben via Stop+Start; SST-Rejoin ~2-3min. Siehe `docs/solutions/bug-fixes/2026-09-29-vm302-zombie-postmigration-galera-rejoin.md`
|
||||||
|
- **QEMU-Cmdline-Falle:** `-incoming unix:/run/qemu-server/NNN.migrate -S` bleibt nach Migration im Prozess-String — KEIN Beweis für "paused". Immer `qm monitor <vmid> <<< 'info status'` prüfen.
|
||||||
|
- **HA-Guard:** `ha-manager disable` existiert nicht. Bei Resource-State `request_stop` geht `qm stop` trotzdem durch.
|
||||||
|
|
||||||
## Related Skills
|
## Related Skills
|
||||||
- `mariadb-galera-cluster-administration` (devops)
|
- `mariadb-galera-cluster-administration` (devops)
|
||||||
|
|
||||||
|
|||||||
+4
-5
@@ -3,7 +3,7 @@ title: "noris AI Platform"
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [ai, llm, noris, gpu, embeddings]
|
tags: [ai, llm, noris, gpu, embeddings]
|
||||||
created: "2026-09-27"
|
created: "2026-09-27"
|
||||||
modified: "2026-09-27"
|
modified: "2026-09-29"
|
||||||
---
|
---
|
||||||
|
|
||||||
# noris AI Platform (ai.noris.de)
|
# noris AI Platform (ai.noris.de)
|
||||||
@@ -13,7 +13,7 @@ modified: "2026-09-27"
|
|||||||
## Endpoints
|
## Endpoints
|
||||||
- **Chat:** `https://ai.noris.de/v1/chat/completions`
|
- **Chat:** `https://ai.noris.de/v1/chat/completions`
|
||||||
- **Embeddings:** `https://ai.noris.de/v1/embeddings`
|
- **Embeddings:** `https://ai.noris.de/v1/embeddings`
|
||||||
- **Images:** `https://ai.noris.de/v1/images/generations` (b64_json)
|
- **Images:** ⚠️ `/v1/images/generations` wird vom Bifrost Gateway **NICHT** unterstützt. Image-Gen-Modelle (qwen-image-2-1) werden über `/v1/chat/completions` angesprochen — das Bild kommt als base64-PNG im `content`-Array zurück (Typ `image_url`, `data:image/png;base64,...`). Siehe `references/vllm-image-generation.md` im Skill `serving-llms-vllm`.
|
||||||
|
|
||||||
## Modelle
|
## Modelle
|
||||||
| Typ | Modell-ID | Hinweise |
|
| Typ | Modell-ID | Hinweise |
|
||||||
@@ -25,15 +25,14 @@ modified: "2026-09-27"
|
|||||||
| Mid-range | `qwen3.8-27b` | |
|
| Mid-range | `qwen3.8-27b` | |
|
||||||
| Fast | `ds-v4-flash` | Low-latency, Paperless OCR |
|
| Fast | `ds-v4-flash` | Low-latency, Paperless OCR |
|
||||||
| Embedding | `harrier` | Vektorembeddings |
|
| Embedding | `harrier` | Vektorembeddings |
|
||||||
| Image Gen | `qwen-image` | via `/v1/images/gen` |
|
| Image Gen | `qwen-image-2-1` | Via `/v1/chat/completions` (NOT images/generations). Base64-PNG im content-Array. ~30s/ Bild. |
|
||||||
| Image Gen | `qsu` | |
|
|
||||||
| Image Gen | `qwen-image-2-1` | |
|
|
||||||
|
|
||||||
## Verbraucher
|
## Verbraucher
|
||||||
- **Hermes Agent** — Primärmodell `glm-5-2` via OpenRouter
|
- **Hermes Agent** — Primärmodell `glm-5-2` via OpenRouter
|
||||||
- **HA LLM Vision** — `gemma-4-31b-it` für Bildanalyse (Frigate Events)
|
- **HA LLM Vision** — `gemma-4-31b-it` für Bildanalyse (Frigate Events)
|
||||||
- **Paperless** — `ds-v4-flash` für OCR/Kategorisierung
|
- **Paperless** — `ds-v4-flash` für OCR/Kategorisierung
|
||||||
- **Personal Coach Bot** — `glm-5-2` via noris direkt
|
- **Personal Coach Bot** — `glm-5-2` via noris direkt
|
||||||
|
- **Dynamic Coach** — `glm-5-2` via noris direkt, `qwen-image-2-1` für Visualisierungen
|
||||||
|
|
||||||
## Related
|
## Related
|
||||||
- [[systems/frigate]] — nutzt noris AI für Event-Klassifizierung
|
- [[systems/frigate]] — nutzt noris AI für Event-Klassifizierung
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ title: Proxmox VE Cluster
|
|||||||
category: systems
|
category: systems
|
||||||
tags: [proxmox, virtualization, lxc, qemu, pve]
|
tags: [proxmox, virtualization, lxc, qemu, pve]
|
||||||
created: "2026-04-28"
|
created: "2026-04-28"
|
||||||
modified: "2026-09-26"
|
modified: "2026-09-29"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Proxmox VE Cluster
|
# Proxmox VE Cluster
|
||||||
@@ -50,7 +50,7 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
|||||||
- Benötigte modprobe.d Config:
|
- Benötigte modprobe.d Config:
|
||||||
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
- `blacklist amdgpu` + `blacklist drm` + `blacklist drm_kms_helper`
|
||||||
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
- `options vfio-pci ids=1002:13c0` + `softdep amdgpu pre: vfio-pci`
|
||||||
- Worker-05 (VM 102) läuft auf ms-a2-2 mit funktionierendem GPU-Passthrough
|
- Worker-05 (VM 102) lief auf ms-a2-2 mit funktionierendem GPU-Passthrough. Seit 2026-09-29 auf proxmox6 (Anti-Collocation mit Worker-01). Prüfen ob GPU-Passthrough auf proxmox6 funktioniert!
|
||||||
|
|
||||||
## Bekannte Probleme
|
## Bekannte Probleme
|
||||||
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
- CT110 kaputte libc — Reparatur ausstehend (still stopped)
|
||||||
@@ -60,21 +60,71 @@ pvesh get /cluster/resources --type vm # Alle VMs/CTs
|
|||||||
|
|
||||||
## HA Rules (PVE 9.2 Rules System)
|
## HA Rules (PVE 9.2 Rules System)
|
||||||
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
Seit 2026-09-28: HA Groups → Rules migriert. Anti-Collocation + Node-Affinity.
|
||||||
|
Seit 2026-09-29: RKE2 CP/Worker + Hermes hinzugefügt.
|
||||||
|
|
||||||
**Resource-Affinity (negative = anti-collocation):**
|
**Resource-Affinity (negative = anti-collocation):**
|
||||||
|
|
||||||
|
*RKE2 Control Plane (etcd-Quorum braucht 2/3):*
|
||||||
|
- `rke2-cp-anti-112-122`: vm:112 ↔ vm:122 — nie auf gleichem Host
|
||||||
|
- `rke2-cp-anti-112-126`: vm:112 ↔ vm:126 — nie auf gleichem Host
|
||||||
|
- `rke2-cp-anti-122-126`: vm:122 ↔ vm:126 — nie auf gleichem Host
|
||||||
|
|
||||||
|
*RKE2 Worker:*
|
||||||
|
- `rke2-worker-anti-102-128`: vm:102 ↔ vm:128 — nie auf gleichem Host
|
||||||
|
- `rke2-worker-anti-102-139`: vm:102 ↔ vm:139 — nie auf gleichem Host
|
||||||
|
- `rke2-worker-anti-128-139`: vm:128 ↔ vm:139 — nie auf gleichem Host
|
||||||
|
|
||||||
|
*Galera:*
|
||||||
- `galera-anti-300-301`: vm:300 ↔ vm:301 — nie auf gleichem Host
|
- `galera-anti-300-301`: vm:300 ↔ vm:301 — nie auf gleichem Host
|
||||||
- `galera-anti-300-302`: vm:300 ↔ vm:302 — nie auf gleichem Host
|
- `galera-anti-300-302`: vm:300 ↔ vm:302 — nie auf gleichem Host
|
||||||
- `galera-anti-301-302`: vm:301 ↔ vm:302 — nie auf gleichem Host
|
- `galera-anti-301-302`: vm:301 ↔ vm:302 — nie auf gleichem Host
|
||||||
|
|
||||||
|
*MaxScale:*
|
||||||
- `maxscale-anti-310-311`: vm:310 ↔ vm:311 — nie auf gleichem Host
|
- `maxscale-anti-310-311`: vm:310 ↔ vm:311 — nie auf gleichem Host
|
||||||
|
|
||||||
**Node-Affinity (non-strict, failover allowed):**
|
**Node-Affinity (non-strict, failover allowed):**
|
||||||
|
|
||||||
|
*RKE2 CP:*
|
||||||
|
- `na-vm112`: vm:112 → proxmox3, proxmox5, proxmox7
|
||||||
|
- `na-vm122`: vm:122 → ms-a2-1, ms-a2-2
|
||||||
|
- `na-vm126`: vm:126 → proxmox4, proxmox5, proxmox6
|
||||||
|
|
||||||
|
*RKE2 Worker:*
|
||||||
|
- `na-vm102`: vm:102 → proxmox5, proxmox6, proxmox7 (NICHT ms-a2-2!)
|
||||||
|
- `na-vm128`: vm:128 → ms-a2-2, proxmox5, proxmox7
|
||||||
|
- `na-vm139`: vm:139 → n5pro, proxmox3, proxmox4
|
||||||
|
|
||||||
|
*Hermes:*
|
||||||
|
- `na-vm230`: vm:230 → n5pro, proxmox5, proxmox6
|
||||||
|
|
||||||
|
*Galera/MaxScale:*
|
||||||
- `na-vm300`: vm:300 → n5pro, proxmox3, proxmox4
|
- `na-vm300`: vm:300 → n5pro, proxmox3, proxmox4
|
||||||
- `na-vm301`: vm:301 → proxmox6, proxmox5, proxmox7
|
- `na-vm301`: vm:301 → proxmox6, proxmox5, proxmox7
|
||||||
- `na-vm302`: vm:302 → ms-a2-2, ms-a2-1
|
- `na-vm302`: vm:302 → ms-a2-2, ms-a2-1
|
||||||
- `na-vm310`: vm:310 → proxmox7, proxmox5, proxmox4
|
- `na-vm310`: vm:310 → proxmox7, proxmox5, proxmox4
|
||||||
- `na-vm311`: vm:311 → ms-a2-1, ms-a2-2
|
- `na-vm311`: vm:311 → ms-a2-1, ms-a2-2
|
||||||
|
|
||||||
> ⚠️ PVE 9.2 Constraint: Resources in resource-affinity rules dürfen keine multi-priority node-affinity haben (gleiche Priorität für alle Nodes erforderlich).
|
**Aktuelle Verteilung (alle Anti-Collocation erfüllt):**
|
||||||
|
| Role | VM | Node |
|
||||||
|
|------|----|------|
|
||||||
|
| RKE2 CP-01 | 112 | proxmox3 |
|
||||||
|
| RKE2 CP-02 | 122 | ms-a2-1 |
|
||||||
|
| RKE2 CP-03 | 126 | proxmox4 |
|
||||||
|
| RKE2 Worker-01 | 128 | proxmox5 |
|
||||||
|
| RKE2 Worker-04 | 139 | n5pro |
|
||||||
|
| RKE2 Worker-05 | 102 | proxmox6 |
|
||||||
|
| Hermes-Agent-01 | 230 | n5pro |
|
||||||
|
| Galera db1 | 300 | n5pro |
|
||||||
|
| Galera db2 | 301 | proxmox6 |
|
||||||
|
| Galera db3 | 302 | ms-a2-2 |
|
||||||
|
| MaxScale-01 | 310 | proxmox7 |
|
||||||
|
| MaxScale-02 | 311 | ms-a2-1 |
|
||||||
|
|
||||||
|
> ⚠️ PVE 9.2 Constraints:
|
||||||
|
> - Resources in resource-affinity rules dürfen keine multi-priority node-affinity haben (gleiche Priorität für alle Nodes erforderlich).
|
||||||
|
> - `ha-manager add` MUSS vor `ha-manager rules add` kommen — sonst "cannot use unmanaged resource".
|
||||||
|
> - Bei gleichzeitigem HA-Add + Anti-Collocation-Violation kann HA-Manager deadlocks (beide VMs auf `migrate` fest). Lösung: eine VM temporär aus HA entfernen, manuell migrieren, dann re-add.
|
||||||
|
> - Online-Migration von VMs mit hohen Memory-Writes (>12GB dirty pages) kann `broken pipe` fehlschlagen. Offline-Migration (stop→migrate→start) als Fallback.
|
||||||
|
|
||||||
## Related
|
## Related
|
||||||
- [[systems/ceph-cluster]]
|
- [[systems/ceph-cluster]]
|
||||||
|
|||||||
@@ -25,9 +25,13 @@ modified: "2026-09-17"
|
|||||||
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
| worker-05 | 10.0.30.65 | Worker (GPU renderD128 ✅ on ms-a2-2) |
|
||||||
|
|
||||||
## Storage
|
## Storage
|
||||||
- **Ceph CSI**: ceph-flash (fast), ceph-hdd (bulk)
|
- **Ceph CSI**: ceph-flash (fast/default), ceph-hdd-replica (bulk), cephfs, cephfs-ssd, ceph-media-ec
|
||||||
|
- **⚠️ StorageClass IaC (2026-09-29):** `ceph-flash` + `ceph-hdd-replica` waren MANUELL erstellt (nicht in Git). Jetzt committed unter `clusters/main/storage/`. Alle RBD SCs MÜSSEN `controller-expand-secret-name/namespace` haben — fehlt dies, schlagen Volume Expansions fehl ("provided secret is empty") und triggern Retry-Loops die API Server überlasten.
|
||||||
|
- **⚠️ PVs snapshotten SC-Parameter zur Provisionierungszeit** — SC fixen reicht NICHT; bestehende PVs brauchen individuellen Patch mit `controllerExpandSecretRef` falls sie ohne erstellt wurden.
|
||||||
|
- **⚠️ StorageClass `parameters` sind IMMUTABLE** — können nicht gepatched werden, müssen gelöscht+neu erstellt werden.
|
||||||
- **⚠️ Snapshot-CRD-Flavor (seit Rebuild 2026-08-01):** `volumesnapshotclasses`-CRD (RKE2-Addon `rke2-snapshot-controller-crd`) hat **FLAT-Schema** — `driver`/`deletionPolicy`/`parameters` auf TOP-LEVEL, `spec:` existiert nicht im Schema (Upstream-CRD wäre nested!). Niemals struktur-intuitiv "reparieren" — CRD-Schema lesen. Live-VSCs: `ceph-rbd-snapclass` (default) + `cephfs-snapclass`
|
- **⚠️ Snapshot-CRD-Flavor (seit Rebuild 2026-08-01):** `volumesnapshotclasses`-CRD (RKE2-Addon `rke2-snapshot-controller-crd`) hat **FLAT-Schema** — `driver`/`deletionPolicy`/`parameters` auf TOP-LEVEL, `spec:` existiert nicht im Schema (Upstream-CRD wäre nested!). Niemals struktur-intuitiv "reparieren" — CRD-Schema lesen. Live-VSCs: `ceph-rbd-snapclass` (default) + `cephfs-snapclass`
|
||||||
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`; Specs **100Gi data / 20Gi WAL** (grow-only — Shrink wird von CNPG-Admission verboten, Git immer nach oben alignieren)
|
- **CNPG PostgreSQL**: `postgres-main` Cluster (3/3 Ready), RW Service `postgres-main-rw.postgres.svc.cluster.local:5432`; Specs **100Gi data / 20Gi WAL** (grow-only — Shrink wird von CNPG-Admission verboten, Git immer nach oben alignieren)
|
||||||
|
- **⚠️ etcd auf Ceph RBD (~20ms WAL fsync auf allen CP-Nodes)** — architektonisches Risiko; unter API Server Write Pressure → gRPC DeadlineExceeded Cascade. Lokales NVMe für etcd data dirs empfohlen.
|
||||||
|
|
||||||
## GitOps
|
## GitOps
|
||||||
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
|
- **ArgoCD**: SSH Deploy Keys (read-only) auf Gitea
|
||||||
|
|||||||
Reference in New Issue
Block a user