Files
mw-pfeddersheim-workstation/docs/tech/crash-analysis-2026-03-17.md
T
ja 54ab23dce9 docs(system): analysis reports and EOD workflows
- Created crash analysis for 2026-03-17 OOM event.
- Documented stability test results and memory tuning effects.
- Added PARA-structured workflows for EOD protocol.
- Updated workstation health report for current 16GB swap status.
2026-03-17 19:02:06 +01:00

53 lines
3.7 KiB
Markdown

# Crash Analysis Report - 2026-03-17
## Incident Overview
- **Timestamp**: 2026-03-17 18:33 (approx. 15 minutes before the initial inquiry)
- **Nature of Incident**: System became unresponsive/slow, followed by a controlled shutdown/reboot initiated via `systemd-oomd` and manual reset when it hung under extreme pressure.
- **Root Cause**: Severe memory pressure (RAM + Swap both > 90%).
## Detailed Findings
### 1. Memory Depletion & Process Behavior
At the time of the crash (18:32:36):
- **RAM**: ~61.8GiB used (66,387,353,600 bytes) / 62.7GiB total (98.6%).
- **Swap**: ~3.6GiB used (3,889,168,384 bytes) / 4GiB total (90.5%).
- **Key Culprit**: The unit `kitty-4851-1.scope` (Kitty terminal) was the primary consumer, reaching a **9.4G memory peak** and 1.2G swap peak.
- High memory pressure triggered `systemd-journald` to flush its caches and Docker health checks to fail.
### 2. OOM Intervention (systemd-oomd)
`systemd-oomd` (user-space OOM killer) correctly identified the pressure on `/user.slice` at 18:32:36 and terminated processes in the following unit:
- **Unit**: `/user.slice/user-1000.slice/user@1000.service/app.slice/kitty-4851-1.scope`
- **Result**: 16 processes were killed within this scope.
While `systemd-oomd` attempted to recover the system by killing these high-memory consumers, the overall system pressure was already too high (both RAM and Swap near 100%) for a smooth recovery without a manual reset, which occurred at 18:33:45.
### 3. Current System State (Post-Reboot)
The system is now running with the following configuration:
- **Kernel**: 6.12.73-1-MANJARO
- **Swap**: 16GB file (`/swapfile`).
- **zswap**: **Active** (zstd/zsmalloc, 20% max pool). This was previously "pending" and is now successfully loaded.
- **systemd-oomd**: Running and monitoring `/user.slice` with 90% swap and 60%/30s memory pressure limits.
- **Hardware**: NVMe disk self-assessment (S.M.A.R.T.) result: **PASSED**.
## Recommendations & Optimizations
### 1. Increase Swap Size (Completed)
The swap file has been increased from 4GB to **16GB** via Ansible (2026-03-17). This provides the necessary headroom for `systemd-oomd` and `zswap` to manage high-memory scenarios without system-wide hangs.
### 2. Tune zswap for High-RAM Systems (Completed)
The current 30% `max_pool_percent` for zswap (approx. 19.2GB) is configured via GRUB and active at runtime. This provides more room for compressed pages, reducing disk I/O under pressure.
### 3. Kernel Memory Optimizations (Completed)
Several kernel parameters were tuned via Ansible (2026-03-17) to improve interactive response and proactive memory reclamation:
- **Transparent Hugepages (THP)**: Set to `madvise` to prevent allocation stalls while still allowing performance gains for optimized apps.
- **Proactive Reclaim**: `vm.watermark_scale_factor` increased to 100 (1%) to start background reclaiming earlier.
- **Dirty Page Management**: `vm.dirty_ratio` and `vm.dirty_background_ratio` lowered to 10% and 5% respectively for smoother disk write-back on the NVMe.
- **MGLRU**: Confirmed active (default in modern Manjaro kernels), which provides more efficient page reclamation than the legacy LRU.
### 4. Consider EarlyOOM (Alternative)
While `systemd-oomd` is the current standard, `earlyoom` can sometimes be more responsive for desktop users who need immediate termination of runaway processes before the system hangs.
- **Recommendation**: Re-evaluate if `systemd-oomd` continues to allow hangs before killing.
## Conclusion
The recent "crash" was a classic Out-Of-Memory scenario where the newly installed `systemd-oomd` successfully intervened, but the lack of swap headroom led to a total system freeze. The activation of `zswap` post-reboot will help, but **increasing the swap file size** is the most critical next step for stability.