GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xuanqi, Shang, Haoyang, Li, Xiaoxiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
di: Huang, Yancheng, et al.
Pubblicazione: (2026)
di: Huang, Yancheng, et al.
Pubblicazione: (2026)
Understanding and Mitigating Memorization in Diffusion Models for Tabular Data
di: Fang, Zhengyu, et al.
Pubblicazione: (2024)
di: Fang, Zhengyu, et al.
Pubblicazione: (2024)
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
di: Bossy, Thierry, et al.
Pubblicazione: (2025)
di: Bossy, Thierry, et al.
Pubblicazione: (2025)
How Do Flow Matching Models Memorize and Generalize in Sample Data Subspaces?
di: Gao, Weiguo, et al.
Pubblicazione: (2024)
di: Gao, Weiguo, et al.
Pubblicazione: (2024)
Memorizing Long-tail Data Can Help Generalization Through Composition
di: Zhou, Mo, et al.
Pubblicazione: (2025)
di: Zhou, Mo, et al.
Pubblicazione: (2025)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
di: Joshi, Kunj, et al.
Pubblicazione: (2025)
di: Joshi, Kunj, et al.
Pubblicazione: (2025)
Localizing and Mitigating Memorization in Image Autoregressive Models
di: Kasliwal, Aditya, et al.
Pubblicazione: (2025)
di: Kasliwal, Aditya, et al.
Pubblicazione: (2025)
Mitigating Memorization In Language Models
di: Sakarvadia, Mansi, et al.
Pubblicazione: (2024)
di: Sakarvadia, Mansi, et al.
Pubblicazione: (2024)
SteerConf: Steering LLMs for Confidence Elicitation
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks
di: Mohan, Vamshi Sunku, et al.
Pubblicazione: (2026)
di: Mohan, Vamshi Sunku, et al.
Pubblicazione: (2026)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
di: Siddique, Zara, et al.
Pubblicazione: (2025)
di: Siddique, Zara, et al.
Pubblicazione: (2025)
Rethinking LoRA for Data Heterogeneous Federated Learning: Subspace and State Alignment
di: Peng, Hongyi, et al.
Pubblicazione: (2026)
di: Peng, Hongyi, et al.
Pubblicazione: (2026)
Memorized Images in Diffusion Models share a Subspace that can be Located and Deleted
di: Chavhan, Ruchika, et al.
Pubblicazione: (2024)
di: Chavhan, Ruchika, et al.
Pubblicazione: (2024)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
di: Ghosh, Shaona, et al.
Pubblicazione: (2025)
di: Ghosh, Shaona, et al.
Pubblicazione: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
di: Li, Aochong Oliver, et al.
Pubblicazione: (2025)
di: Li, Aochong Oliver, et al.
Pubblicazione: (2025)
TARDIS: Mitigating Temporal Misalignment via Representation Steering
di: Shin, Changho, et al.
Pubblicazione: (2025)
di: Shin, Changho, et al.
Pubblicazione: (2025)
Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
di: Jeon, Dongjae, et al.
Pubblicazione: (2024)
di: Jeon, Dongjae, et al.
Pubblicazione: (2024)
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
di: Gadgil, Soham, et al.
Pubblicazione: (2026)
di: Gadgil, Soham, et al.
Pubblicazione: (2026)
Mitigating Memorization in LLMs using Activation Steering
di: Suri, Manan, et al.
Pubblicazione: (2025)
di: Suri, Manan, et al.
Pubblicazione: (2025)
Conjunction Subspaces Test for Conformal and Selective Classification
di: He, Zengyou, et al.
Pubblicazione: (2024)
di: He, Zengyou, et al.
Pubblicazione: (2024)
Steering LLMs via Scalable Interactive Oversight
di: Zhou, Enyu, et al.
Pubblicazione: (2026)
di: Zhou, Enyu, et al.
Pubblicazione: (2026)
FineGates: LLMs Finetuning with Compression using Stochastic Gates
di: Svirsky, Jonathan, et al.
Pubblicazione: (2024)
di: Svirsky, Jonathan, et al.
Pubblicazione: (2024)
(Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs
di: Tao, Jiashu, et al.
Pubblicazione: (2025)
di: Tao, Jiashu, et al.
Pubblicazione: (2025)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
di: Yu, Ziming, et al.
Pubblicazione: (2024)
di: Yu, Ziming, et al.
Pubblicazione: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts
di: Busch, Matthias, et al.
Pubblicazione: (2026)
di: Busch, Matthias, et al.
Pubblicazione: (2026)
Memorize Early, Then Query: Inlier-Memorization-Guided Active Outlier Detection
di: Kang, Minseo, et al.
Pubblicazione: (2026)
di: Kang, Minseo, et al.
Pubblicazione: (2026)
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
di: Karnik, Sathwik, et al.
Pubblicazione: (2026)
di: Karnik, Sathwik, et al.
Pubblicazione: (2026)
The Pitfalls of Memorization: When Memorization Hurts Generalization
di: Bayat, Reza, et al.
Pubblicazione: (2024)
di: Bayat, Reza, et al.
Pubblicazione: (2024)
MemHunter: Automated and Verifiable Memorization Detection at Dataset-scale in LLMs
di: Wu, Zhenpeng, et al.
Pubblicazione: (2024)
di: Wu, Zhenpeng, et al.
Pubblicazione: (2024)
Gated Subspace Inference for Transformer Acceleration
di: Thomas, Stephen J.
Pubblicazione: (2026)
di: Thomas, Stephen J.
Pubblicazione: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
di: Li, Lujun, et al.
Pubblicazione: (2025)
di: Li, Lujun, et al.
Pubblicazione: (2025)
Selective Safety Steering via Value-Filtered Decoding
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2026)
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2026)
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Mitigating Cognitive Inertia in Large Reasoning Models via Latent Spike Steering
di: Lee, Seojin, et al.
Pubblicazione: (2026)
di: Lee, Seojin, et al.
Pubblicazione: (2026)
Understanding and Mitigating Dataset Corruption in LLM Steering
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
di: Vargas, Francisco, et al.
Pubblicazione: (2020)
di: Vargas, Francisco, et al.
Pubblicazione: (2020)
DSO: Direct Steering Optimization for Bias Mitigation
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025) -
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
di: Huang, Yancheng, et al.
Pubblicazione: (2026) -
Understanding and Mitigating Memorization in Diffusion Models for Tabular Data
di: Fang, Zhengyu, et al.
Pubblicazione: (2024) -
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
di: Bossy, Thierry, et al.
Pubblicazione: (2025) -
How Do Flow Matching Models Memorize and Generalize in Sample Data Subspaces?
di: Gao, Weiguo, et al.
Pubblicazione: (2024)