Science AI
Preprint describes long-context jailbreak technique
A preprint introduces NINJA, a jailbreak method that appends benign model-generated content to harmful goals in long language-model contexts. The authors report that harmful-goal position affects safety performance and that NINJA increased attack success rates on the HarmBench benchmark across LLaMA, Qwen, Mistral and Gemini models. They also report that, under a fixed compute budget, longer context can outperform more best-of-N trials.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire


