Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghilardi, Davide, Belotti, Federico, Molinari, Marco, Ma, Tao, Palmonari, Matteo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating LLMs on Entity Disambiguation in Tables
von: Belotti, Federico, et al.
Veröffentlicht: (2024)
von: Belotti, Federico, et al.
Veröffentlicht: (2024)
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
von: Bono, Carlo, et al.
Veröffentlicht: (2025)
von: Bono, Carlo, et al.
Veröffentlicht: (2025)
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
von: Chen, Boshui, et al.
Veröffentlicht: (2026)
von: Chen, Boshui, et al.
Veröffentlicht: (2026)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
von: Joshi, Ananya, et al.
Veröffentlicht: (2025)
von: Joshi, Ananya, et al.
Veröffentlicht: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
Pruning Large Language Models with Semi-Structural Adaptive Sparse Training
von: Huang, Weiyu, et al.
Veröffentlicht: (2024)
von: Huang, Weiyu, et al.
Veröffentlicht: (2024)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
Arithmetic with Language Models: from Memorization to Computation
von: Maltoni, Davide, et al.
Veröffentlicht: (2023)
von: Maltoni, Davide, et al.
Veröffentlicht: (2023)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
von: Maltoni, Davide, et al.
Veröffentlicht: (2025)
von: Maltoni, Davide, et al.
Veröffentlicht: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models
von: Huber, Bernd, et al.
Veröffentlicht: (2025)
von: Huber, Bernd, et al.
Veröffentlicht: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
A Study on Hidden Layer Distillation for Large Language Model Pre-Training
von: Guigon, Maxime, et al.
Veröffentlicht: (2026)
von: Guigon, Maxime, et al.
Veröffentlicht: (2026)
ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
von: Pozzi, Riccardo, et al.
Veröffentlicht: (2025)
von: Pozzi, Riccardo, et al.
Veröffentlicht: (2025)
Constrain Alignment with Sparse Autoencoders
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
Large Language Models for Automatic Milestone Detection in Group Discussions
von: Duan, Zhuoxu, et al.
Veröffentlicht: (2024)
von: Duan, Zhuoxu, et al.
Veröffentlicht: (2024)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
von: Fang, Yi, et al.
Veröffentlicht: (2026)
von: Fang, Yi, et al.
Veröffentlicht: (2026)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
In-context Autoencoder for Context Compression in a Large Language Model
von: Ge, Tao, et al.
Veröffentlicht: (2023)
von: Ge, Tao, et al.
Veröffentlicht: (2023)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
von: Zhong, Chengzhi, et al.
Veröffentlicht: (2025)
von: Zhong, Chengzhi, et al.
Veröffentlicht: (2025)
Reversing Large Language Models for Efficient Training and Fine-Tuning
von: Gal, Eshed, et al.
Veröffentlicht: (2025)
von: Gal, Eshed, et al.
Veröffentlicht: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
Evolving Subnetwork Training for Large Language Models
von: Li, Hanqi, et al.
Veröffentlicht: (2024)
von: Li, Hanqi, et al.
Veröffentlicht: (2024)
Sparse Autoencoders for Hypothesis Generation
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
Group Preference Optimization: Few-Shot Alignment of Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2023)
von: Zhao, Siyan, et al.
Veröffentlicht: (2023)
Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
LaCo: Large Language Model Pruning via Layer Collapse
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
Leveraging Adaptive Group Negotiation for Heterogeneous Multi-Robot Collaboration with Large Language Models
von: Song, Siqi, et al.
Veröffentlicht: (2025)
von: Song, Siqi, et al.
Veröffentlicht: (2025)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating LLMs on Entity Disambiguation in Tables
von: Belotti, Federico, et al.
Veröffentlicht: (2024) -
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
von: Bono, Carlo, et al.
Veröffentlicht: (2025) -
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
von: Liu, Xiangyu, et al.
Veröffentlicht: (2025) -
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
von: Chen, Boshui, et al.
Veröffentlicht: (2026) -
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)