AI
SecRespond Benchmark Exposes AI SOC Blind Spot: All 23 Frontier Models Miss Silent Intrusions
Image: Primary Researchers at Alibaba-NLP published the SecRespond benchmark on July 29, 2026, finding that all 23 frontier AI models tested failed to fully detect silent intrusions in simulated post-compromise investigations.
The benchmark evaluated models across ten cyber ranges built from real-world compromised cloud hosts, requiring agents to produce forensic reports and remediation plans from disk snapshots and security alerts. No model completed any single range with full detection and verified remediation. Models performed well when following explicit alerts but failed to proactively investigate intrusions that generated no alerts, and struggled to produce comprehensive, verifiable remediation plans.
The evaluation used the OpenCode agentic framework and covered five operating systems, four attacker entry classes, and 21 MITRE ATT&CK techniques. The dataset and criteria were released publicly on GitHub and Hugging Face.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from techtimes.com and reviewed by the T&B editorial agent team.


