Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Xuansheng, Yuan, Jiayi, Yao, Wenlin, Zhai, Xiaoming, Liu, Ninghao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
di: Wang, Yun, et al.
Pubblicazione: (2025)
di: Wang, Yun, et al.
Pubblicazione: (2025)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
di: Härle, Ruben, et al.
Pubblicazione: (2024)
di: Härle, Ruben, et al.
Pubblicazione: (2024)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2026)
di: Cho, Seonglae, et al.
Pubblicazione: (2026)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
di: Guo, Shuchen, et al.
Pubblicazione: (2025)
di: Guo, Shuchen, et al.
Pubblicazione: (2025)
Investigating CoT Monitorability in Large Reasoning Models
di: Yang, Shu, et al.
Pubblicazione: (2025)
di: Yang, Shu, et al.
Pubblicazione: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
di: Fang, Yi, et al.
Pubblicazione: (2026)
di: Fang, Yi, et al.
Pubblicazione: (2026)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
di: Ghussin, Yusser Al, et al.
Pubblicazione: (2026)
di: Ghussin, Yusser Al, et al.
Pubblicazione: (2026)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
A Comparative Analysis of Sparse Autoencoder and Activation Difference in Language Model Steering
di: Xie, Jiaqing
Pubblicazione: (2025)
di: Xie, Jiaqing
Pubblicazione: (2025)
Foundation Models for Low-Resource Language Education (Vision Paper)
di: Ding, Zhaojun, et al.
Pubblicazione: (2024)
di: Ding, Zhaojun, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of ASR models using Sparse Autoencoders
di: Pluth, Dan, et al.
Pubblicazione: (2026)
di: Pluth, Dan, et al.
Pubblicazione: (2026)
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
di: Du, Hongfei, et al.
Pubblicazione: (2026)
di: Du, Hongfei, et al.
Pubblicazione: (2026)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
Interpretable Company Similarity with Sparse Autoencoders
di: Molinari, Marco, et al.
Pubblicazione: (2024)
di: Molinari, Marco, et al.
Pubblicazione: (2024)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
di: Kurochkin, Vadim, et al.
Pubblicazione: (2025)
di: Kurochkin, Vadim, et al.
Pubblicazione: (2025)
VOLTA: Improving Generative Diversity by Variational Mutual Information Maximizing Autoencoder
di: Ma, Yueen, et al.
Pubblicazione: (2023)
di: Ma, Yueen, et al.
Pubblicazione: (2023)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
di: Goyal, Agam, et al.
Pubblicazione: (2025)
di: Goyal, Agam, et al.
Pubblicazione: (2025)
InFoBench: Evaluating Instruction Following Ability in Large Language Models
di: Qin, Yiwei, et al.
Pubblicazione: (2024)
di: Qin, Yiwei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025) -
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025) -
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
di: Wang, Yun, et al.
Pubblicazione: (2026) -
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)