Skip to main content
Back to Newswire
Science AI

arXiv paper proposes statistically grounded SAE steering for LLM activation control

An arXiv preprint on artificial intelligence presents a transparent sparse autoencoder feature steering pipeline for activation-space control of large language models, framed as a lightweight alternative to fine-tuning. The method first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three statistics: an F-test, KSG mutual information, and Cohen's d. The steering direction is built as a Cohen's-d-weighted combination of SAE decoder rows, described as an optimization-free construction motivated by Fisher-LDA under approximate feature decorrelation. Across three Gemma-family models, four behavioral domains, and 356 layer-strength configurations, the approach produced measurable domain-specific shifts while showing a substantial gap between raw attribute movement and quality-preserving generation. In the strongest configuration reported, logical-correctness steering reached a primary-score delta of +1.16 in Gemma 2 9B. The authors find that usable steering is highly localized by model, domain, layer, and strength, and argue that activation-steering evaluations should report quality-conditioned success alongside raw behavioral shift. Code and data are stated to be available with the paper.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.