Skip to main content
Back to Newswire
Science AI

arXiv paper proposes statistically grounded SAE steering for LLM activation control

An arXiv preprint on artificial intelligence presents a transparent sparse autoencoder feature steering pipeline for activation-space control of large language models, framed as a lightweight alternative to fine-tuning. The method first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three statistics: an F-test, KSG mutual information, and Cohen's d. The steering direction is built as a Cohen's-d-weighted combination of SAE decoder rows, described as an optimization-free construction motivated by Fisher-LDA under approximate feature decorrelation. Across three Gemma-family models, four behavioral domains, and 356 layer-strength configurations, the approach produced measurable domain-specific shifts while showing a substantial gap between raw attribute movement and quality-preserving generation. In the strongest configuration reported, logical-correctness steering reached a primary-score delta of +1.16 in Gemma 2 9B. The authors find that usable steering is highly localized by model, domain, layer, and strength, and argue that activation-steering evaluations should report quality-conditioned success alongside raw behavioral shift. Code and data are stated to be available with the paper.
Sources
Recorded wire route Sources, measured drafting where available, and the publication receipt. See concurrent Machine
Evidence entered
Admission Evidence and chronology passed Science AI
Publication receipt Entered the validated Newswire
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.
Back to Newswire