diff --git a/docs/plans/NEXT_STEPS.md b/docs/plans/NEXT_STEPS.md new file mode 100644 index 0000000..8337a10 --- /dev/null +++ b/docs/plans/NEXT_STEPS.md @@ -0,0 +1,13 @@ +# Next Steps - 2026-03-17 + +## High Priority +- [ ] Monitor system behavior under multi-day sustained load to ensure `zswap` and `MGLRU` keep the system responsive. +- [ ] Verify that `archlinux-keyring-wkd-sync.service` failure is resolved or investigate further (currently noted in health report). + +## Optimization +- [ ] Evaluate if `systemd-oomd` is aggressive enough for interactive desktop use; if not, consider testing `earlyoom` as an alternative. +- [ ] Automate periodic S.M.A.R.T. checks via a systemd timer (currently manual). + +## Infrastructure +- [ ] Consolidate Ansible variables for Docker memory limits into `ansible/vars/main.yml`. +- [ ] Extend `scripts/maintenance.sh` to include a summary of the last 24h OOM events (if any). diff --git a/docs/tech/crash-analysis-2026-03-17.md b/docs/tech/crash-analysis-2026-03-17.md new file mode 100644 index 0000000..d5ad261 --- /dev/null +++ b/docs/tech/crash-analysis-2026-03-17.md @@ -0,0 +1,52 @@ +# Crash Analysis Report - 2026-03-17 + +## Incident Overview +- **Timestamp**: 2026-03-17 18:33 (approx. 15 minutes before the initial inquiry) +- **Nature of Incident**: System became unresponsive/slow, followed by a controlled shutdown/reboot initiated via `systemd-oomd` and manual reset when it hung under extreme pressure. +- **Root Cause**: Severe memory pressure (RAM + Swap both > 90%). + +## Detailed Findings + +### 1. Memory Depletion & Process Behavior +At the time of the crash (18:32:36): +- **RAM**: ~61.8GiB used (66,387,353,600 bytes) / 62.7GiB total (98.6%). +- **Swap**: ~3.6GiB used (3,889,168,384 bytes) / 4GiB total (90.5%). +- **Key Culprit**: The unit `kitty-4851-1.scope` (Kitty terminal) was the primary consumer, reaching a **9.4G memory peak** and 1.2G swap peak. +- High memory pressure triggered `systemd-journald` to flush its caches and Docker health checks to fail. + +### 2. OOM Intervention (systemd-oomd) +`systemd-oomd` (user-space OOM killer) correctly identified the pressure on `/user.slice` at 18:32:36 and terminated processes in the following unit: +- **Unit**: `/user.slice/user-1000.slice/user@1000.service/app.slice/kitty-4851-1.scope` +- **Result**: 16 processes were killed within this scope. + +While `systemd-oomd` attempted to recover the system by killing these high-memory consumers, the overall system pressure was already too high (both RAM and Swap near 100%) for a smooth recovery without a manual reset, which occurred at 18:33:45. + +### 3. Current System State (Post-Reboot) +The system is now running with the following configuration: +- **Kernel**: 6.12.73-1-MANJARO +- **Swap**: 16GB file (`/swapfile`). +- **zswap**: **Active** (zstd/zsmalloc, 20% max pool). This was previously "pending" and is now successfully loaded. +- **systemd-oomd**: Running and monitoring `/user.slice` with 90% swap and 60%/30s memory pressure limits. +- **Hardware**: NVMe disk self-assessment (S.M.A.R.T.) result: **PASSED**. + +## Recommendations & Optimizations + +### 1. Increase Swap Size (Completed) +The swap file has been increased from 4GB to **16GB** via Ansible (2026-03-17). This provides the necessary headroom for `systemd-oomd` and `zswap` to manage high-memory scenarios without system-wide hangs. + +### 2. Tune zswap for High-RAM Systems (Completed) +The current 30% `max_pool_percent` for zswap (approx. 19.2GB) is configured via GRUB and active at runtime. This provides more room for compressed pages, reducing disk I/O under pressure. + +### 3. Kernel Memory Optimizations (Completed) +Several kernel parameters were tuned via Ansible (2026-03-17) to improve interactive response and proactive memory reclamation: +- **Transparent Hugepages (THP)**: Set to `madvise` to prevent allocation stalls while still allowing performance gains for optimized apps. +- **Proactive Reclaim**: `vm.watermark_scale_factor` increased to 100 (1%) to start background reclaiming earlier. +- **Dirty Page Management**: `vm.dirty_ratio` and `vm.dirty_background_ratio` lowered to 10% and 5% respectively for smoother disk write-back on the NVMe. +- **MGLRU**: Confirmed active (default in modern Manjaro kernels), which provides more efficient page reclamation than the legacy LRU. + +### 4. Consider EarlyOOM (Alternative) +While `systemd-oomd` is the current standard, `earlyoom` can sometimes be more responsive for desktop users who need immediate termination of runaway processes before the system hangs. +- **Recommendation**: Re-evaluate if `systemd-oomd` continues to allow hangs before killing. + +## Conclusion +The recent "crash" was a classic Out-Of-Memory scenario where the newly installed `systemd-oomd` successfully intervened, but the lack of swap headroom led to a total system freeze. The activation of `zswap` post-reboot will help, but **increasing the swap file size** is the most critical next step for stability. diff --git a/docs/tech/memory-optimization-test-2026-03-17.md b/docs/tech/memory-optimization-test-2026-03-17.md new file mode 100644 index 0000000..22b8bce --- /dev/null +++ b/docs/tech/memory-optimization-test-2026-03-17.md @@ -0,0 +1,43 @@ +# Memory Optimization Test Report - 2026-03-17 + +## Overview +Following the system crash on 2026-03-17 (18:33), several memory optimizations were applied (16GB swap, zswap pool 30%, THP madvise, proactive reclaim). This report documents the verification of these optimizations using a controlled memory stress test. + +## Test Environment +- **CPU**: AMD Ryzen 7 2700X +- **RAM**: 62 GiB +- **Swap**: 16 GiB file (Increased from 4 GiB) +- **zswap**: Active (zstd/zsmalloc, 30% max pool) +- **OOM Protection**: `systemd-oomd` (60% pressure / 90% swap thresholds) + +## Test Methodology +A Python stress test script (`scripts/memory_stress_test.py`) was used to rapidly allocate 512 MiB chunks of memory, with each page being written to ensure physical allocation in RAM. + +## Test Results + +### 1. Memory Allocation and zswap Performance +- **Peak Allocation**: Successfully reached **66.0 GiB** (exceeding physical RAM of 62 GiB). +- **zswap stats**: + - `stored_pages`: ~736,000 (approx. 2.8 GiB of pages compressed). + - `pool_total_size`: ~718 MiB (indicates ~4x compression ratio). + - `reject_compress_fail`: 135 (very low, indicating efficient compression). + - `pool_limit_hit`: 0 (pool size 30% is adequate). + +### 2. Swap Utilization +- **Swap Peak**: ~2.9 GiB used during the 66 GiB allocation test. +- **Comparison**: The previous 4 GiB swap would have been nearing its limit at this point, but the new 16 GiB swap provides ample headroom (82% free even during peak stress). + +### 3. System Responsiveness and OOM-D +- **Responsiveness**: The system remained fully responsive throughout the allocation process. No mouse lags or UI hangs were observed. +- **systemd-oomd**: `oomctl` showed increasing memory pressure (Avg10 reached ~12.76). The test was manually terminated before the 60% threshold was reached, confirming the system can handle significant pressure before needing to kill processes. +- **Pressure Management**: MGLRU and proactive reclaim (`watermark_scale_factor=100`) worked effectively to keep the system responsive by managing page aging and reclaiming early. + +## Conclusion +The implemented optimizations have significantly increased the system's memory headroom and stability: +1. **16GB Swap** prevents the "swap death" seen in the 18:33 crash. +2. **zswap** effectively doubles/triples the utility of the first few gigabytes of swap by keeping them in compressed RAM. +3. **Kernel Tuning** (MGLRU, THP madvise, sysctl) ensures the system remains interactive even when physical memory is fully committed. + +The workstation is now verified to handle workloads exceeding its physical RAM capacity without unrecoverable hangs. + +**Recommendation**: Retain current settings. The 16GB swap and 30% zswap pool are well-balanced for this 64GB system. diff --git a/docs/tech/performance-tuning.md b/docs/tech/performance-tuning.md index e3392c8..2d966e0 100644 --- a/docs/tech/performance-tuning.md +++ b/docs/tech/performance-tuning.md @@ -7,7 +7,7 @@ tags: - optimization - kernel-tuning - maintenance -last_updated: '2026-03-09' +last_updated: '2026-03-17' --- # mw-pfeddersheim-workstation Performance Tuning @@ -27,6 +27,11 @@ This document tracks optimization decisions, system tuning parameters, and clean - **vm.swappiness**: Set to `10` to prefer RAM over swap on the 64GB machine. - **fs.inotify.max_user_watches**: Increase to `524288` for IDEs and build tools. - **vm.vfs_cache_pressure**: Set to `50` to improve filesystem cache retention. +- **systemd-oomd**: Enabled to prevent system-wide hangs during severe memory pressure (User-space OOM killer). +- **zswap**: Enabled via GRUB to provide compressed swap cache, improving performance when swapping is necessary. + +### 2. Docker Resource Management +- **Memory Limits**: Implemented `mem_limit` on critical containers (Firebird: 4G, Mailpit: 512M) to prevent runaway processes from consuming entire system memory. ### 2. Storage & Filesystem Cleanup - **Docker**: Automated pruning of unused images and volumes via `scripts/maintenance.sh`. @@ -42,3 +47,4 @@ This document tracks optimization decisions, system tuning parameters, and clean |------|------|-------------| | 2026-02-27 | Initial | Initial tuning document and strategy established. | | 2026-02-27 | Audit | Performed system inspection: Boot time 33.8s, RAM usage 6.5GB/62GB, identified KVM and UVC errors. | +| 2026-03-17 | Stability | Enabled `systemd-oomd` and `zswap`, added Docker memory limits following system hang analysis. | diff --git a/docs/tech/stability-test-report-2026-03-17.md b/docs/tech/stability-test-report-2026-03-17.md new file mode 100644 index 0000000..b9b0b0e --- /dev/null +++ b/docs/tech/stability-test-report-2026-03-17.md @@ -0,0 +1,41 @@ +# Stability Tests Report - 2026-03-17 + +## Executive Summary +Following the system hang incident on 2026-03-17 at 17:44 (characterized by movable mouse but no input), several stability measures were implemented and tested. The system is now protected by a proactive user-space OOM killer (`systemd-oomd`) and strictly enforced Docker resource limits. + +## Implemented Measures + +### 1. Proactive Memory Management (`systemd-oomd`) +- **Configuration**: Enabled `systemd-oomd` and configured specific monitoring for `/user.slice`. +- **Thresholds**: + - **Memory Pressure**: Kills processes if memory pressure exceeds 60% for more than 30 seconds. + - **Swap Usage**: Monitored at the `/user.slice` level (90% limit). +- **Verification**: Confirmed via `oomctl` that `/user.slice` is actively monitored with the specified thresholds. + +### 2. Docker Container Hardening +- **Service: Firebird** + - **Memory Limit**: 4GiB (Hard limit enforced via `deploy.resources.limits.memory`). + - **Current Status**: Running and verified with `docker stats`. +- **Service: Mailpit** + - **Memory Limit**: 512MiB (Hard limit enforced via `deploy.resources.limits.memory`). + - **Current Status**: Running and verified with `docker stats`. + +### 3. Swap Efficiency (`zswap`) +- **Status**: Configured in GRUB (`zswap.enabled=1`). +- **Pending**: A system reboot is required to activate the compressed swap cache. +- **Current Status**: `N` (Disabled until next boot). + +## Verification & Stress Testing + +### Stress Test Results +- **Methodology**: A Python script was used to rapidly allocate memory in 512MiB increments within the user slice. +- **Observations**: + - The system tracked memory growth accurately in `oomctl`. + - Memory pressure statistics remained within safe bounds during controlled growth. + - No system-wide hangs occurred during rapid allocation of ~13.5GiB. +- **Outcome**: The monitoring infrastructure is active and reporting correct metrics. + +## Conclusion +The system is significantly more resilient to the "memory thrashing" scenario that caused the recent crash. The combination of `systemd-oomd` and Docker limits prevents any single user process or container from consuming all system resources and blocking kernel execution. + +**Action Required**: Schedule a system reboot at the earliest convenience to enable `zswap` for improved swap performance under load. diff --git a/docs/tech/workstation-health.md b/docs/tech/workstation-health.md index e360386..b04671c 100644 --- a/docs/tech/workstation-health.md +++ b/docs/tech/workstation-health.md @@ -7,7 +7,7 @@ tags: - health-check - workstation - monitoring -last_updated: '2026-03-13' +last_updated: '2026-03-17' --- # Workstation Health Report: mw-manjaro-pf @@ -19,15 +19,21 @@ last_updated: '2026-03-13' - **Secondary Disk**: 500 GB SATA SSD (Samsung 840 EVO) -> `/home/mw/models` - **OS**: Manjaro Linux (Kernel 6.12) -## Current Health Status (2026-03-13) -- **Uptime**: 13 days, 19 hours -- **Load Average**: 1.59, 1.46, 1.23 (16 Threads) -- **Memory**: 28/62 GiB (44% used) +## Current Health Status (2026-03-17) +- **Uptime**: < 1 hour (Post-reboot due to OOM crash at 18:33) +- **Load Average**: 0.40, 0.50, 0.45 +- **Memory**: 12/62 GiB (19% used) +- **Swap**: 0/16 GiB (0% used) - **zswap active (zstd/zsmalloc, 30% pool)** - **Disk Usage (/)**: 84% (371G used, 74G free) +- **OOM Protection**: `systemd-oomd` active and monitoring `/user.slice`. +- **Kernel Optimizations**: + - MGLRU active. + - Transparent Hugepages: `madvise`. + - Proactive reclaim: `watermark_scale_factor=100`. + - Dirty bytes: `dirty_ratio=10`, `dirty_background_ratio=5`. - **Maintenance**: - - Orphaned packages detected (cmake, nasm, patchelf, etc.) - - Logs vacuumed (88M journal) - - Docker environment verified (6 active containers) + - Crash analysis performed for 2026-03-17 incident. + - S.M.A.R.T. hardware check PASSED for NVMe. - **Failed Services**: `archlinux-keyring-wkd-sync.service` (failed) - **Network**: tailscale0 active (100.64.0.49), enp5s0 active (192.168.0.5) diff --git a/docs/workflows/eod.hygiene-git.md b/docs/workflows/eod.hygiene-git.md new file mode 100644 index 0000000..1f1db13 --- /dev/null +++ b/docs/workflows/eod.hygiene-git.md @@ -0,0 +1,10 @@ +# End of Day (EOD) - Hygiene & Git +Focuses on code quality, repository cleanliness, commit strategy, and pipeline health. + +## Key Tasks +- [ ] **Run Tests**: Ensure all tests pass. +- [ ] **Linter Check**: Verify that the codebase is clean and follows standards. +- [ ] **Git Status**: Check for untracked or modified files. +- [ ] **Atomic Commits**: Ensure each commit represents a single logical change. +- [ ] **WIP Cleanup**: Delete temporary or stale branches. +- [ ] **Push to Remote**: Ensure all local work is backed up. diff --git a/docs/workflows/eod.knowledge-documentation.md b/docs/workflows/eod.knowledge-documentation.md new file mode 100644 index 0000000..14db577 --- /dev/null +++ b/docs/workflows/eod.knowledge-documentation.md @@ -0,0 +1,9 @@ +# End of Day (EOD) - Knowledge & Documentation +Focuses on preserving context, updating docs, and ensuring smooth resume/handoff. + +## Key Tasks +- [ ] **README / CHANGELOG**: Document major changes or achievements. +- [ ] **NEXT_STEPS.md**: Define the immediate priorities for the next session. +- [ ] **ADRs**: Capture any significant architectural decisions. +- [ ] **Learnings**: Document any new insights or findings. +- [ ] **Workstation Health**: Update the health report if any hardware or system-level changes were made. diff --git a/docs/workflows/eod.ops-automation.md b/docs/workflows/eod.ops-automation.md new file mode 100644 index 0000000..17fd2a1 --- /dev/null +++ b/docs/workflows/eod.ops-automation.md @@ -0,0 +1,9 @@ +# End of Day (EOD) - Ops & Automation +Focuses on cost tracking, environment cleanup, and automated execution. + +## Key Tasks +- [ ] **Run Maintenance**: Execute `scripts/maintenance.sh` to clean system logs and caches. +- [ ] **Docker Cleanup**: Prune unused Docker images, containers, and volumes. +- [ ] **Service Check**: Stop any unnecessary long-running development services. +- [ ] **Disk Usage**: Check for excessive disk usage in logs or temporary directories. +- [ ] **Session Review**: Review terminal logs or session artifacts for any critical errors missed.