Hello Ahmad Ibrahim,
Based on the behavior you described, audit event loss during periods of high system activity is commonly associated with the audit backlog queue filling faster than auditd can process records. As a best practice, I recommend increasing the backlog_limit gradually rather than making a large change immediately, while monitoring backlog utilization, CPU consumption, memory usage, and audit event generation rates to validate the impact.
To determine an appropriate value, start by measuring the peak audit event volume during your busiest workload periods and compare it against the rate at which auditd is able to process and write events. Systems with higher CPU and available memory can generally sustain larger backlog queues, but the optimal setting depends on the specific audit policy and workload characteristics. In addition to backlog_limit, I suggest reviewing related parameters such as backlog_wait_time, rate_limit, flush, freq, and the failure handling options configured in auditd.conf and the audit ruleset.
It is also important to verify that excessive or redundant audit rules are not generating unnecessary events, as rule optimization can significantly reduce queue pressure. Monitoring kernel audit messages for backlog warnings, dropped events, and queue saturation indicators can help identify whether the bottleneck is event generation, disk I/O, or auditd processing performance. After any tuning changes, I recommend conducting controlled load testing to confirm that the backlog remains within acceptable limits and that no audit records are being dropped under peak conditions.
I hope the response provided some helpful insight. If you find this answer useful, please hit “accept answer” so I know it addressed your concern.
Jason