SOC-AgentBench Leaderboard
Aria AI Security Research — Community benchmark for AI security analyst agents.
Compare models on technique detection, evidence quality, timeline accuracy, and safety metrics.
Dataset: alirezaaminzadeh/soc-agent-bench
| Split | Episodes |
|---|---|
| train | 84 |
| validation | 18 |
| test | 18 |
Benchmark Results (Test Split)
soc-analyst-tool-use (QLoRA) | fine-tuned | 0.7222222222222221 | 0.78 | 0.71 | 0.65 | 0.68 | 0.92 | 0.09 | 0.04 | 18.4 |
Metrics
| Metric | Description |
|---|---|
| Technique Macro F1 | MITRE ATT&CK technique detection F1 |
| Evidence Precision | Fraction of cited evidence that is valid |
| Evidence Recall | Fraction of gold evidence cited |
| Root Cause Accuracy | Correct incident root cause identification |
| Timeline Accuracy | Event timeline overlap F1 |
| Tool Success Rate | Successful tool invocations |
| Hallucination Rate | Fabricated evidence citations |
| Unsafe Action Rate | Destructive actions without evidence |
Submit Results
To submit benchmark results, open a Discussion on the
soc-agent-bench dataset
with your eval_results.json and model ID.