Modern Security Operations Centers face growing volumes of SIEM alerts, increasing the risk of analyst burnout, lower triage quality, and slower incident response. This thesis presents WARDEN, an autonomous AI agent for first-level investigation of alerts generated by Wazuh, an open-source SIEM based on OpenSearch. WARDEN follows a multi-phase architecture: a deterministic pre-filter based on known false positives, an LLM validator for borderline cases, and a lead investigator agent that performs evidence-driven analysis through playbooks. For high-severity alerts or low-confidence outcomes, a challenger reviews the verdict before it is stored. The system was evaluated on 60 real alerts collected in production. The deterministic phase resolved 15% of cases in 0.1 seconds on average without token consumption. All 60 investigations required 233 seconds on average and reached a 90% analyst satisfaction rate based on dashboard feedback. The verdict distribution reflects a conservative design: since false negatives can have serious consequences, escalation to human review is treated as a valid outcome. The results show that agent-based AI can support SOC automation while preserving analyst oversight. The evaluation also highlights limitations: reliance on external LLM providers, limited context awareness in the false-positive knowledge base, and occasional investigation scope expansion. These aspects guide future work, starting with on-premises language model deployment.
WARDEN: Design and Evaluation of an Agentic AI System for Autonomous SOC Alert Triage
ZANE, FILIPPO
2025/2026
Abstract
Modern Security Operations Centers face growing volumes of SIEM alerts, increasing the risk of analyst burnout, lower triage quality, and slower incident response. This thesis presents WARDEN, an autonomous AI agent for first-level investigation of alerts generated by Wazuh, an open-source SIEM based on OpenSearch. WARDEN follows a multi-phase architecture: a deterministic pre-filter based on known false positives, an LLM validator for borderline cases, and a lead investigator agent that performs evidence-driven analysis through playbooks. For high-severity alerts or low-confidence outcomes, a challenger reviews the verdict before it is stored. The system was evaluated on 60 real alerts collected in production. The deterministic phase resolved 15% of cases in 0.1 seconds on average without token consumption. All 60 investigations required 233 seconds on average and reached a 90% analyst satisfaction rate based on dashboard feedback. The verdict distribution reflects a conservative design: since false negatives can have serious consequences, escalation to human review is treated as a valid outcome. The results show that agent-based AI can support SOC automation while preserving analyst oversight. The evaluation also highlights limitations: reliance on external LLM providers, limited context awareness in the false-positive knowledge base, and occasional investigation scope expansion. These aspects guide future work, starting with on-premises language model deployment.| File | Dimensione | Formato | |
|---|---|---|---|
|
ZANE_FILIPPO_880119.pdf
accesso aperto
Dimensione
1.98 MB
Formato
Adobe PDF
|
1.98 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14247/29704