- Move 12 historical/superseded docs + 1 session log + 1 PDF + 1 HTML plan to archive/ - Merge 4 VM-crash docs into archive/vm-crash-incident.md - Merge LDAP plan + spec into ldap.md - Merge FilePond race investigation into filepond-crash-analysis.md - Merge SPECS.md client notes into spec-sheet.md appendix - Update README index and security.md cross-reference
5.4 KiB
VM Crash Root Cause Analysis (posterg.erg.be)
Date: 2026-03-26 Server: posterg.erg.be Status: ✅ ROOT CAUSE IDENTIFIED — NOT the application's fault
Merged from
VM_Crash_Analysis_FINAL.md,VM_Crash_Reports.md,EVIDENCE_SUMMARY.md, andIMMEDIATE_FIX.md(single incident).
🔥 ROOT CAUSE: Serial Console (serial-getty) Crash Loop
The VM did not crash due to the nginx/posterg application. The crash was
caused by a systemd serial-getty@ttyS0 service crash loop that ran
continuously for ~50 days, eventually exhausting system memory.
The smoking gun
- 1,264,488 serial-getty crashes recorded in the journal
- Restart counter reached 421,491 by the time of the OOM event
- Crashed every 10 seconds for the entire uptime
- Error:
agetty[PID]: could not get terminal name: -22/failed to get terminal attributes: Input/output error
Timeline reconstruction
| Date | Event | Details |
|---|---|---|
| Jan 13, 2026 | System boot | Clean boot, services started normally |
| Jan 13 – Mar 4 | Serial getty crash loop | ~421,491 restarts over 48.7 days (6 restarts/min) |
| Mar 4, 10:45 | MariaDB memory pressure | InnoDB reports memory pressure event |
| Mar 4, 10:50 | OOM Killer triggered | Systemd invokes OOM killer due to memory exhaustion |
| Mar 4, 10:51 | Journal stops | System likely became unresponsive |
| Mar 4 – Mar 24 | Unknown state | 20-day gap in logs |
| Mar 24, 12:56 | Hard reboot | Technicians forced reboot |
| Mar 24, 12:57 | System back online | New boot, clean state |
Why this happened
QEMU/KVM virtual machine configuration issue. The error
could not get terminal name: -22 (EINVAL) indicates the VM's serial console
(ttyS0) is misconfigured or not properly connected at the hypervisor level.
Common causes: serial console enabled in VM config but not attached to host,
QEMU -serial parameter misconfigured, VirtIO console driver issue, or
host-side serial device permissions.
Resource impact
Each agetty spawn creates a process, opens file descriptors, and logs to the
journal (~200 bytes per entry). Over 50 days at 6 crashes/minute:
- ~421,000 failed process spawns
- ~1.2 million journal entries (~240MB journal bloat)
- Gradual memory exhaustion → OOM killer
🔍 The application is NOT at fault
Evidence the posterg application is innocent:
- No PHP-FPM crashes — clean operation, 11.1–11.2M peak memory
- No nginx errors before OOM — the 234KB error log is from after the reboot (Mar 26), mostly blocked security-scanner attempts
- Normal traffic — only internal IP 192.168.6.11 accessing the site
- No DB issues before crash — SQLite working fine
Post-reboot issues (unrelated to crash)
After the Mar 24 reboot there were schema errors (no such table: tags,
no such column: ts.role) caused by code updates (Mar 24 14:49) without a
matching migration — not the crash cause.
Post-reboot security events (Mar 26)
955 blocked requests from 192.168.6.11 (.env, .git/config, WordPress/
Next.js/Nuxt.js probes) — all properly blocked by nginx (working as designed).
🛠️ The fix
Disable the broken serial console service:
sudo systemctl stop serial-getty@ttyS0.service
sudo systemctl disable serial-getty@ttyS0.service
sudo systemctl mask serial-getty@ttyS0.service
# Verify
sudo systemctl status serial-getty@ttyS0.service # → "Loaded: masked"
Also fix the post-reboot DB schema errors:
cd /var/www/posterg
ls -la storage/migrations/
sqlite3 storage/posterg.db "SELECT name FROM sqlite_master WHERE type='table';"
Optional: fix the serial console properly (hypervisor)
If serial console access is needed for emergency recovery, configure it on the
QEMU/KVM host via virsh edit posterg (add/verify <serial type='pty'> +
<console ...>), restart the VM in a maintenance window, then unmask/re-enable
serial-getty@ttyS0.
📊 Post-reboot system health
✅ All systems healthy — memory 6% used, disk 12% used, swap unused, load idle. nginx 4 workers, PHP-FPM 2 workers, MariaDB 155MB RSS (all normal).
🎯 Recommendations
- CRITICAL: disable
serial-getty@ttyS0(see fix above) - Fix DB schema for post-reboot errors
- Improve monitoring —
prometheus-node-exporteror systemd unit monitoring would have surfaced the serial-getty loop earlier - Journal maintenance:
sudo journalctl --disk-usage sudo journalctl --vacuum-size=500M sudo journalctl --vacuum-time=30d # /etc/systemd/journald.conf: SystemMaxUse=500M, SystemKeepFree=1G, MaxRetentionSec=30day - Optional: tighten
limit_reqrates and addfail2banfor repeated 403s
📎 Appendix: technical details
OOM event
Mar 04 10:50:23 posterg kernel: systemd invoked oom-killer
gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0
Serial getty error code
agetty[PID]: could not get terminal name: -22 — EINVAL, terminal
initialization on a misconfigured ttyS0 device.
Journal statistics
- Total journal entries: ~193 MB
- Serial-getty crashes: 1,264,488 (~65% of journal)
- Uptime at OOM: ~50 days (Jan 13 – Mar 4)
- Crash frequency: every 10s; total restarts 421,491
Report prepared by: Automated analysis + human review Confidence: 🟢 HIGH (definitively identified from kernel/journal/service logs) Risk: before fix 🟠 HIGH (will recur ~50 days) · after fix 🟢 LOW