docs: update workstation health with crashlog analysis
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
---
|
||||
description: >-
|
||||
Comprehensive health report for the mw-pfeddersheim-workstation. including
|
||||
Comprehensive health report for the mw-pfeddersheim-workstation, including
|
||||
system specifications, current status, and runtime environment details.
|
||||
tags:
|
||||
- infrastructure
|
||||
@@ -15,60 +15,76 @@ last_updated: '2026-03-17'
|
||||
|
||||
- **CPU**: AMD Ryzen 7 2700X (16 threads) @ 3.9GHz max
|
||||
- **Memory**: 62 GiB DDR5-3200 MHz
|
||||
- **Primary Disk**: 512 GB NVMe SSD (Samsung MZvl2512HCJQ,)
|
||||
- **Primary Disk**: 512 GB NVMe SSD (Samsung MZvl2512HCJQ)
|
||||
- **Secondary Disk**: 500 GB SATA SSD (Samsung 840 EVO) -> `/home/mw/models`
|
||||
- **OS**: Manjaro Linux (Kernel 6.12)
|
||||
## Current Health Status (2026-03-17)
|
||||
- **uptime**: <1 hour (post-reboot due to Oom crash at 18:33)
|
||||
- **Load Average**: 0.40, 0.50, 1.45
|
||||
- **Memory**: 12/62 GiB (19% used)
|
||||
- **Swap**: 0/16GiB (0% used) - **zswap active**: Yes (zstd/zsmalloc, 30% pool)
|
||||
|
||||
## Current Health Status (2026-03-17 23:21 CET)
|
||||
|
||||
- **Uptime**: ~1 hour, 7 minutes (post-reboot due to OOM crash at 18:33)
|
||||
- **Load Average**: 0.40, 0.50, 1.45 (very low)
|
||||
- **Memory**: 7.9 GiB used / 62 GiB total (12%)
|
||||
- **Swap**: 0 B used / 16 GiB (0%)
|
||||
- **zswap**: Active (zstd/zsmalloc, 30% pool)
|
||||
- **Disk Usage (/)**: 84% (371G used, 74G free)
|
||||
- **OOM Protection**: `systemd-oomd` active and monitoring `/user.slice`.
|
||||
- **Kernel Optimizations**:
|
||||
- mglru active.
|
||||
- **OOM Protection**: `systemd-oomd` active, monitoring `/user.slice`
|
||||
- **Kernel Optimizations**:
|
||||
- mglru: Active
|
||||
- Transparent hugepages: `madvise`
|
||||
- Proactive reclaim: `watermark_scale_factor 100`.
|
||||
- Proactive reclaim: `watermark_scale_factor 100`
|
||||
- Dirty bytes: `dirty_ratio=10`, `dirty_background_ratio=5`
|
||||
- **Maintenance**:
|
||||
- Crash analysis performed for 2026-03-17 incident.
|
||||
- S.M.A.R.T. hardware check passed for NVMe.
|
||||
**Failed services**: `archlinux-keyring-wkd-sync.service` (failed)
|
||||
**Network**: tailscale (present but dormant)
|
||||
- `archlinux-timesync` is not active since 2024, but news about issues
|
||||
- The user may look into them `wkd-sync` service status (`systemctl status wkd-sync --no-pager`) shows the is dead
|
||||
- However, `systemctl` doesn't report status for units it `wkd-sync` expects `systemd` unit but output (which may be spurious).
|
||||
- **Maintenance**:
|
||||
- Crash analysis performed for 2026-03-17 incident
|
||||
- S.M.A.R.T. hardware check passed for NVMe
|
||||
|
||||
- **Dell iDRivet**: Check battery (CR20322 - next check scheduled)
|
||||
- **Fstrim**: File system trimming (weekly)
|
||||
- **Docker containers**: Check for zombie containers (`docker ps | grep -a 'zombie' | awk '{print $2}' | head -1; done`)
|
||||
## Crashlog Analysis (2026-03-17)
|
||||
|
||||
# Check for zombie containers
|
||||
# Skip docker daemon socket file to avoid docker zombies
|
||||
ps aux | head -1 | grep -a 'CONT.*' | grep -a zombie
|
||||
for container in "${docker_container_name}"
|
||||
if ! systemctl is-active "${container_name}"; then
|
||||
# Check: 0 for dead container processes
|
||||
for container in "${docker_container_name}" in $(docker ps | grep -a 'zombie' | awk '{print $2}')
|
||||
done
|
||||
fi
|
||||
done
|
||||
### Recent Errors (from journalctl)
|
||||
|
||||
- **Crash analysis**: Previous Oom crash analyzed. critical context.
|
||||
System is stable now with optimizations applied.
|
||||
No new issues detected.
|
||||
</code>
|
||||
No errors in the last hour from `journalctl -p err`
|
||||
|
||||
### OOM Events
|
||||
|
||||
- **15:06:06 CET**: Node.js process (PID 122570, 122573, 122574, 122575, 122625) dumped core due to OOM
|
||||
- These appear to be Cline/agent-related processes that exhausted memory
|
||||
- Systemd-oomd correctly intervened and terminated the processes
|
||||
|
||||
### Kernel Warnings (Boot)
|
||||
|
||||
- **VMSCAPE**: SMT advisory - mitigation available (enable STIBP for full protection)
|
||||
- **ATA link failures**: ata3/4/7/8 failed to resume (idle SATA ports)
|
||||
- **Samsung 840 EVO quirks**: noncqtrim, zeroaftertrim, nodmalog (normal for older SSD)
|
||||
- **NVIDIA taint**: Proprietary driver (expected)
|
||||
- **UVC webcam warnings**: Non-critical (device firmware)
|
||||
- **Pipewire Broken Pipe**: Video device initialization issues (non-critical)
|
||||
|
||||
### Failed Services
|
||||
|
||||
- `archlinux-keyring-wkd-sync.service`: Failed (non-critical, keyring sync)
|
||||
|
||||
### Docker
|
||||
|
||||
- **Status**: 6 containers running
|
||||
- **No zombie containers detected**
|
||||
|
||||
### Health Check Script Results
|
||||
|
||||
```text
|
||||
Root partition: 78% used
|
||||
Git status: Dirty (34 changes) - workstation-health.md modified
|
||||
Docker: 6 containers running
|
||||
Uptime: ~1 hour, 7 minutes
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
1. Consider setting up automated crash monitoring for high-memory applications.
|
||||
2 - **Replace Samsung 840 EVO** - currently applies SATA quirks. TRIM and NCQ. Can be disabled, by the is a of NVMe performance on modern kernels.
|
||||
- Action: No action needed (non-critical)
|
||||
## Recommendations
|
||||
|
||||
- **Monitor memory**: 12% usage (52GiB total, 7% used) is very healthy.
|
||||
- **GPU**: 12GB usage (54GiB available) 55GiB free)
|
||||
- **Run morning health check script** (`/home/mw/internal/mw-pfeddersheim-workstation/scripts/morning-health-check.sh`).
|
||||
- Check existing docs
|
||||
- [ ] Run morning health check script
|
||||
</task_progress>
|
||||
</write_to_file>
|
||||
1. **Monitor Node.js memory**: Today's OOM crash involved multiple Cline-related node processes.
|
||||
Consider investigating if agent sessions are consuming excessive memory.
|
||||
2. **Replace Samsung 840 EVO**: Currently applies SATA quirks (noncqtrim, zeroaftertrim, nodmalog).
|
||||
Action: No action needed (non-critical, working as expected)
|
||||
3. **Monitor memory**: 12% usage is very healthy. Current state is stable.
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. Consider setting up automated crash monitoring for high-memory applications
|
||||
2. Run morning health check script daily: `/home/mw/internal/mw-pfeddersheim-workstation/scripts/morning-health-check.sh`
|
||||
Reference in New Issue
Block a user