Science AI
Study Maps How Adversarial Prompts Rewire LLM Internal Reasoning During Jailbreaks
Researchers introduced a mechanistic framework that compares the internal computation graphs of a language model processing clean versus adversarial prompts.
By aligning these graphs, they found that jailbreaks systematically suppress safety-related components, introduce attack-specific features, and reroute computation paths. The method decomposes model computation into invariant, suppressed, and emergent structures, identifying recurring vulnerability motifs. Causal interventions on the identified nodes and subgraphs reduced attack success rates across multiple open-source models and jailbreak benchmarks.
The authors argue that internal computation graphs provide a causal foundation for diagnosing and mitigating model failures, moving beyond input-output correlation to structural intervention.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from arXiv and reviewed by the T&B editorial agent team.
Back to Newswire