SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Dengcan, Li, Jiahao, Fu, Zheren, Tu, Yi, Li, Jiajun, Mao, Zhendong, Zhang, Yongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Constrain Alignment with Sparse Autoencoders
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA
von: Li, Jiaang, et al.
Veröffentlicht: (2024)
von: Li, Jiaang, et al.
Veröffentlicht: (2024)
On-the-fly Preference Alignment via Principle-Guided Decoding
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
von: Fang, Yi, et al.
Veröffentlicht: (2026)
von: Fang, Yi, et al.
Veröffentlicht: (2026)
DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
von: Wang, Shaohan, et al.
Veröffentlicht: (2025)
von: Wang, Shaohan, et al.
Veröffentlicht: (2025)
Sparse Autoencoders for Hypothesis Generation
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
von: Yang, Xianjun, et al.
Veröffentlicht: (2025)
von: Yang, Xianjun, et al.
Veröffentlicht: (2025)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2025)
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2024)
von: Chanin, David, et al.
Veröffentlicht: (2024)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
von: Joshi, Ananya, et al.
Veröffentlicht: (2025)
von: Joshi, Ananya, et al.
Veröffentlicht: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
ToolRM: Towards Agentic Tool-Use Reward Modeling
von: Li, Renhao, et al.
Veröffentlicht: (2025)
von: Li, Renhao, et al.
Veröffentlicht: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
von: Muchane, Mark, et al.
Veröffentlicht: (2025)
von: Muchane, Mark, et al.
Veröffentlicht: (2025)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
von: Ghilardi, Davide, et al.
Veröffentlicht: (2024)
von: Ghilardi, Davide, et al.
Veröffentlicht: (2024)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
von: Li, T. Ed, et al.
Veröffentlicht: (2025)
von: Li, T. Ed, et al.
Veröffentlicht: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
von: Zhang, Ruikang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruikang, et al.
Veröffentlicht: (2026)
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
von: Xia, Hou, et al.
Veröffentlicht: (2025)
von: Xia, Hou, et al.
Veröffentlicht: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
Towards Understanding the Robustness of Sparse Autoencoders
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
PrunePath: Towards Highly Structured Sparse Language Models
von: Gu, Zhexuan, et al.
Veröffentlicht: (2026)
von: Gu, Zhexuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Constrain Alignment with Sparse Autoencoders
von: Yin, Qingyu, et al.
Veröffentlicht: (2024) -
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2025) -
Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
von: Zhu, Mingye, et al.
Veröffentlicht: (2025) -
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
von: Shi, Wei, et al.
Veröffentlicht: (2025) -
ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA
von: Li, Jiaang, et al.
Veröffentlicht: (2024)