Skip to main content

Share story

Science AI

Study Maps How Adversarial Prompts Rewire LLM Internal Reasoning During Jailbreaks

Researchers introduced a mechanistic framework that compares the internal computation graphs of a language model processing clean versus adversarial prompts. By aligning these graphs, they found that jailbreaks systematically suppress safety-related components, introduce attack-specific features, and reroute computation paths. The method decomposes model computation into invariant, suppressed, and emergent structures, identifying recurring vulnerability motifs. Causal interventions on the identified nodes and subgraphs reduced attack success rates across multiple open-source models and jailbreak benchmarks. The authors argue that internal computation graphs provide a causal foundation for diagnosing and mitigating model failures, moving beyond input-output correlation to structural intervention.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Science
AI Science

Preprint reports schema-free generation of valid enterprise test data

Researchers report in an arXiv preprint that their Generalist Populator agent generated enterprise data with 100% constraint satisfaction and 0.88 average marginal fidelity across ten simulated environments without accessing datab...

AI Science
AI Science

Preprint reports task-completion gains from agent-generated interfaces

Researchers report in an arXiv preprint that training an agent to generate interactive interfaces improved a 4B model's Pass@3 task-completion score from 9.33% to 58.00%. Their GenUI-Harness pairs an agent that retrieves informati...

Robotics AI
Robotics AI

DreamTrue researchers report fewer interaction defects in robot video predictions

Researchers report that DreamTrue, a model that predicts videos of robot actions, reduced human-assessed interaction defects from 48.12% to 6.25% on AgiBot. The preprint addresses predictions that follow commands inaccurately or f...

Robotics AI
Robotics AI

FAITH preprint reports humanoid safety gains while preserving task performance

Researchers report that their FAITH safety filter achieved a 99.95% safety rate while retaining 97% of unfiltered task return on a 29-degree-of-freedom humanoid in Walking-Avoid. The preprint also describes demonstrations of the s...

AI Science
AI Science

AMD, OpenAI and Google join AI resource pledges for federal science

AMD, OpenAI and Google are pledging AI tools and compute credits for the federal Genesis Mission alongside Nvidia and Anthropic, The Next Web reported. AMD's commitment is $500M, OpenAI's is $200M and Google's is $150M. The pledg...