SOC-AgentBench Leaderboard

Aria AI Security Research — Community benchmark for AI security analyst agents.

Compare models on technique detection, evidence quality, timeline accuracy, and safety metrics.

Dataset: alirezaaminzadeh/soc-agent-bench

Split Episodes
train 84
validation 18
test 18

Benchmark Results (Test Split)

Metrics

Metric Description
Technique Macro F1 MITRE ATT&CK technique detection F1
Evidence Precision Fraction of cited evidence that is valid
Evidence Recall Fraction of gold evidence cited
Root Cause Accuracy Correct incident root cause identification
Timeline Accuracy Event timeline overlap F1
Tool Success Rate Successful tool invocations
Hallucination Rate Fabricated evidence citations
Unsafe Action Rate Destructive actions without evidence

Submit Results

To submit benchmark results, open a Discussion on the soc-agent-bench dataset with your eval_results.json and model ID.

Related