Skip to main content

Share story

Science AI

HARC Couples Harmfulness and Refusal Directions for Stronger LLM Safety Alignment

Researchers introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs harmfulness and refusal directions across both prompt and response positions in the residual stream. The method achieves the strongest robustness-capability-usability trade-off among six baselines spanning major training-time and inference-time safety methods. Prior work showed aligned LLMs encode harmfulness and refusal as separable directions at prompt-side token positions. The new analysis extends to response-token positions, finding models recognize harmful content while generating it, even when failing to recognize the input as harmful at the prompt side. Jailbreaks succeed by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Since HARC confines intervention to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact without degrading general capability or inflating over-refusal. The harmfulness and refusal directions at prompt and response positions transfer across five model families and two scales without architecture-specific tuning.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Products
Products

Paramount closes $110 billion Warner Bros. Discovery acquisition

Paramount has completed its $110 billion acquisition of Warner Bros. Discovery, creating a combined company called Skydance, The Verge reports. The transaction brings the companies' film studios, streaming services and brands incl...

AI Capital
AI Capital

DeepSeek reportedly nears funding of at least 80 billion yuan

DeepSeek is close to raising at least 80 billion yuan in a new funding round, Bloomberg reported Tuesday, with Tencent and CATL committing some of the largest sums. The AI lab initially sought about 50 billion yuan; demand increas...

Capital AI
Capital AI

Vinci raises $250M to expand physics simulation software

Vinci announced Tuesday that it raised $250M at a $1.5bn valuation to expand software that predicts how chip and hardware designs behave physically before they are built. Advent, Temasek and Xora Innovation led the round, with Ecl...

Capital AI
Capital AI

Flai raises $27 million for dealership AI software

Flai said Tuesday it raised a $27 million Series A led by Base10 Partners for its AI software serving car dealerships. Investors included Friedkin Group, Findlay Automotive, Toyota's venture arm, Y Combinator and First Round Capit...

Security Capital
Security Capital

Hadrian raises $40m to expand automated security testing

Amsterdam cybersecurity startup Hadrian has raised $40m in a round co-led by Forgepoint Capital International and Smartfin, bringing its total funding to $65m. The company says it will expand across Europe, the Middle East, Africa...

Capital
Capital

Vinci raises $250 million at $1.5 billion valuation

Software startup Vinci said Tuesday it raised $250 million at a $1.5 billion valuation, Reuters reported. The company makes software that simulates elements of chip and other hardware design and is seeking to expand its suite of s...