Tests find older Claude models bypass explicit-content safeguards
Image: Primary
Image: Primary AWS says its Continuum security system passed 819 of 920 tasks on CyberGym-E2E within the benchmark's 90-minute limit, achieving an 89.0% success rate. That exceeds the previous public high of 65.9% by 23.1 percentage points. The...
Independent researcher Syed Anas Mohiuddin demonstrated attacks that use one internal AI agent to pass malicious instructions to other agents, Ars Technica reported. Google and four other organizations have acknowledged related vu...
Google, JPMorgan Chase, Weaviate and two government teams have fixed flaws that could let attackers direct AI tool servers toward internal systems, The Next Web reported. Researcher Syed Anas Mohiuddin reported all five vulnerabil...
The Wikimedia Foundation disclosed that its investigation found unauthorized wiki edits, unsuccessful attempts to exploit its public note-taking tool and heavy traffic from agents it attributes to OpenAI. It found no evidence that...
CoreWeave says Agent Lens is available in public preview within Forge. The tool reads agents' OpenTelemetry records and groups conversations with similar failures, linking each group to the underlying execution traces. Developers...
GitHub released a research preview of ReviewBench, a public benchmark that lets developers evaluate AI code reviewers against a common dataset and scoring method. It contains 219 pull requests from 187 public repositories spanning...