Modern Security Operations Centers face growing volumes of SIEM alerts, increasing the risk of analyst burnout, lower triage quality, and slower incident response. This thesis presents WARDEN, an autonomous AI agent for first-level investigation of alerts generated by Wazuh, an open-source SIEM based on OpenSearch. WARDEN follows a multi-phase architecture: a deterministic pre-filter based on known false positives, an LLM validator for borderline cases, and a lead investigator agent that performs evidence-driven analysis through playbooks. For high-severity alerts or low-confidence outcomes, a challenger reviews the verdict before it is stored. The system was evaluated on 60 real alerts collected in production. The deterministic phase resolved 15% of cases in 0.1 seconds on average without token consumption. All 60 investigations required 233 seconds on average and reached a 90% analyst satisfaction rate based on dashboard feedback. The verdict distribution reflects a conservative design: since false negatives can have serious consequences, escalation to human review is treated as a valid outcome. The results show that agent-based AI can support SOC automation while preserving analyst oversight. The evaluation also highlights limitations: reliance on external LLM providers, limited context awareness in the false-positive knowledge base, and occasional investigation scope expansion. These aspects guide future work, starting with on-premises language model deployment.

WARDEN: Design and Evaluation of an Agentic AI System for Autonomous SOC Alert Triage

ZANE, FILIPPO
2025/2026

Abstract

Modern Security Operations Centers face growing volumes of SIEM alerts, increasing the risk of analyst burnout, lower triage quality, and slower incident response. This thesis presents WARDEN, an autonomous AI agent for first-level investigation of alerts generated by Wazuh, an open-source SIEM based on OpenSearch. WARDEN follows a multi-phase architecture: a deterministic pre-filter based on known false positives, an LLM validator for borderline cases, and a lead investigator agent that performs evidence-driven analysis through playbooks. For high-severity alerts or low-confidence outcomes, a challenger reviews the verdict before it is stored. The system was evaluated on 60 real alerts collected in production. The deterministic phase resolved 15% of cases in 0.1 seconds on average without token consumption. All 60 investigations required 233 seconds on average and reached a 90% analyst satisfaction rate based on dashboard feedback. The verdict distribution reflects a conservative design: since false negatives can have serious consequences, escalation to human review is treated as a valid outcome. The results show that agent-based AI can support SOC automation while preserving analyst oversight. The evaluation also highlights limitations: reliance on external LLM providers, limited context awareness in the false-positive knowledge base, and occasional investigation scope expansion. These aspects guide future work, starting with on-premises language model deployment.
File in questo prodotto:
File Dimensione Formato  
ZANE_FILIPPO_880119.pdf

accesso aperto

Dimensione 1.98 MB
Formato Adobe PDF
1.98 MB Adobe PDF Visualizza/Apri

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14247/29704