Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
Computer Science > Artificial Intelligence arXiv:2607.19364 (cs) [Submitted on 5 Jun 2026] Title:Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models Authors:Oshayer Siddique, J. We introduce a transparent SAE-feature steering pipeline that first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three complementary statistics: $F$-test, KSG mutual information, and Cohen's $d$. The resulting steering direction is constructed as a Cohen's-$d$-weighted combination of SAE decoder rows, providing an optimization-free direction motivated by Fisher-LDA under approximate SAE-feature decorrelation.
- ▪Computer Science > Artificial Intelligence arXiv:2607.19364 (cs) [Submitted on 5 Jun 2026] Title:Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models Authors:Oshayer Siddique, J.
- ▪We introduce a transparent SAE-feature steering pipeline that first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three complementary statistics: $F$-test, KSG mutual infor
- ▪The resulting steering direction is constructed as a Cohen's-$d$-weighted combination of SAE decoder rows, providing an optimization-free direction motivated by Fisher-LDA under approximate SAE-feature decorrelation.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Artificial Intelligence arXiv:2607.19364 (cs) [Submitted on 5 Jun 2026] Title:Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models Authors:Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan View a PDF of the paper titled Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models, by Oshayer Siddique and 5 other authors View PDF HTML (experimental) Abstract:Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.