Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Dianyun, Ma, Qingsen, Shang, Yuhu, Lu, Zhifeng, Xu, Zhenbo, Ning, Lechen, Wu, Huijia, He, Zhaofeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
by: Ma, Qingsen, et al.
Published: (2025)
by: Ma, Qingsen, et al.
Published: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
Beyond Darkness: Thermal-Supervised 3D Gaussian Splatting for Low-Light Novel View Synthesis
by: Ma, Qingsen, et al.
Published: (2025)
by: Ma, Qingsen, et al.
Published: (2025)
CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
by: Ma, Qingsen, et al.
Published: (2025)
by: Ma, Qingsen, et al.
Published: (2025)
RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation
by: Lu, Zhifeng, et al.
Published: (2026)
by: Lu, Zhifeng, et al.
Published: (2026)
Regularizing Subspace Redundancy of Low-Rank Adaptation
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
Mixture-of-Subspaces in Low-Rank Adaptation
by: Wu, Taiqiang, et al.
Published: (2024)
by: Wu, Taiqiang, et al.
Published: (2024)
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
by: Zhou, Guanghao, et al.
Published: (2026)
by: Zhou, Guanghao, et al.
Published: (2026)
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
by: Li, Mingjie, et al.
Published: (2025)
by: Li, Mingjie, et al.
Published: (2025)
A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
by: Xu, Xiaoang, et al.
Published: (2025)
by: Xu, Xiaoang, et al.
Published: (2025)
AI Safety, Alignment, and Ethics (AI SAE)
by: Waldner, Dylan
Published: (2025)
by: Waldner, Dylan
Published: (2025)
Subspace Geometry Governs Catastrophic Forgetting in Low-Rank Adaptation
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
SAE-V: Interpreting Multimodal Models for Enhanced Alignment
by: Lou, Hantao, et al.
Published: (2025)
by: Lou, Hantao, et al.
Published: (2025)
Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language Models
by: Lu, Yuheng, et al.
Published: (2024)
by: Lu, Yuheng, et al.
Published: (2024)
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
by: Kong, Zicheng, et al.
Published: (2026)
by: Kong, Zicheng, et al.
Published: (2026)
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Rethinking Class-Incremental Learning from a Dynamic Imbalanced Learning Perspective
by: Wang, Leyuan, et al.
Published: (2024)
by: Wang, Leyuan, et al.
Published: (2024)
Subspace Alignment for Vision-Language Model Test-time Adaptation
by: Zeng, Zhichen, et al.
Published: (2026)
by: Zeng, Zhichen, et al.
Published: (2026)
Complementary Subspace Low-Rank Adaptation of Vision-Language Models for Few-Shot Classification
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
SRLoRA: Subspace Recomposition in Low-Rank Adaptation via Importance-Based Fusion and Reinitialization
by: Yang, Haodong, et al.
Published: (2025)
by: Yang, Haodong, et al.
Published: (2025)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace Inference
by: Samplawski, Colin, et al.
Published: (2025)
by: Samplawski, Colin, et al.
Published: (2025)
A Bayesian Interpretation of Adaptive Low-Rank Adaptation
by: Chen, Haolin, et al.
Published: (2024)
by: Chen, Haolin, et al.
Published: (2024)
Unobservable Subspace Evolution and Alignment for Consistent Visual-Inertial Navigation
by: Tian, Chungeng, et al.
Published: (2025)
by: Tian, Chungeng, et al.
Published: (2025)
CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
by: Xiao, Xiaojun, et al.
Published: (2024)
by: Xiao, Xiaojun, et al.
Published: (2024)
Low Rank Adaptation for Adversarial Perturbation
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
Test-time Adaptation for Regression by Subspace Alignment
by: Adachi, Kazuki, et al.
Published: (2024)
by: Adachi, Kazuki, et al.
Published: (2024)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
RaSA: Rank-Sharing Low-Rank Adaptation
by: He, Zhiwei, et al.
Published: (2025)
by: He, Zhiwei, et al.
Published: (2025)
Low-Rank Approximation, Adaptation, and Other Tales
by: Lu, Jun
Published: (2024)
by: Lu, Jun
Published: (2024)
VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
by: Shen, Shufan, et al.
Published: (2025)
by: Shen, Shufan, et al.
Published: (2025)
Low-Rank Compression of Pretrained Models via Randomized Subspace Iteration
by: Pourkamali-Anaraki, Farhad
Published: (2026)
by: Pourkamali-Anaraki, Farhad
Published: (2026)
GLoRIA: Gated Low-Rank Interpretable Adaptation for Dialectal ASR
by: Mehralian, Pouya, et al.
Published: (2026)
by: Mehralian, Pouya, et al.
Published: (2026)
Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
Flat-LoRA: Low-Rank Adaptation over a Flat Loss Landscape
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
CUDA-Accelerated Soft Robot Neural Evolution with Large Language Model Supervision
by: Zhang, Lechen
Published: (2024)
by: Zhang, Lechen
Published: (2024)
SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
by: Zhou, Junchao, et al.
Published: (2025)
by: Zhou, Junchao, et al.
Published: (2025)
Dynamic Generation of Personalities with Large Language Models
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
Similar Items
-
Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
by: Ma, Qingsen, et al.
Published: (2025) -
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026) -
Beyond Darkness: Thermal-Supervised 3D Gaussian Splatting for Low-Light Novel View Synthesis
by: Ma, Qingsen, et al.
Published: (2025) -
CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
by: Ma, Qingsen, et al.
Published: (2025) -
RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation
by: Lu, Zhifeng, et al.
Published: (2026)