An hourly QRadar crash, traced to expired internal certificates
The problem
A QRadar deployment kept crashing on a schedule: roughly one minute of service outage, every hour. For a SIEM, every one of those minutes is a gap in event collection.
Root cause and fix
I traced the instability to expired internal certificates and renewed the internal certificate chain. The hourly outage window disappeared and event collection stabilized.
A separate improvement: reporting
Separately, I configured the QRadar mail server and automated scheduled daily reports to stakeholders, which removed manual report generation from daily operations.
What I took from it
A failure that keeps perfect time is a clue in itself. When something breaks on a schedule, it is worth looking at what else runs on that schedule, and at the plumbing such as certificates, before blaming the application.