Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qi, Wang, Yifei, Cui, Jingyi, Pan, Xiang, Lei, Qi, Jegelka, Stefanie, Wang, Yisen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
An Augmentation Overlap Theory of Contrastive Learning
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Understanding the Role of Equivariance in Self-supervised Learning
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
by: Zhang, Yi-Ge, et al.
Published: (2025)
by: Zhang, Yi-Ge, et al.
Published: (2025)
Non-negative Contrastive Learning
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective
by: Yan, Hanqi, et al.
Published: (2024)
by: Yan, Hanqi, et al.
Published: (2024)
Dissecting the Failure of Invariant Learning on Graphs
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
Can In-context Learning Really Generalize to Out-of-distribution Tasks?
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Do Generated Data Always Help Contrastive Learning?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
by: Guo, Lixuan, et al.
Published: (2026)
by: Guo, Lixuan, et al.
Published: (2026)
Learning from Emergence: A Study on Proactively Inhibiting the Monosemantic Neurons of Artificial Neural Networks
by: Wang, Jiachuan, et al.
Published: (2023)
by: Wang, Jiachuan, et al.
Published: (2023)
A Canonicalization Perspective on Invariant and Equivariant Learning
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
Geometric Algorithms for Neural Combinatorial Optimization with Constraints
by: Karalias, Nikolaos, et al.
Published: (2025)
by: Karalias, Nikolaos, et al.
Published: (2025)
Learning with Exact Invariances in Polynomial Time
by: Soleymani, Ashkan, et al.
Published: (2025)
by: Soleymani, Ashkan, et al.
Published: (2025)
Survey on Generalization Theory for Graph Neural Networks
by: Vasileiou, Antonis, et al.
Published: (2025)
by: Vasileiou, Antonis, et al.
Published: (2025)
Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
by: Phan, Hoang, et al.
Published: (2025)
by: Phan, Hoang, et al.
Published: (2025)
Extracting Interaction-Aware Monosemantic Concepts in Recommender Systems
by: Arviv, Dor, et al.
Published: (2025)
by: Arviv, Dor, et al.
Published: (2025)
On the Stability of Expressive Positional Encodings for Graphs
by: Huang, Yinan, et al.
Published: (2023)
by: Huang, Yinan, et al.
Published: (2023)
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023)
by: Wei, Zeming, et al.
Published: (2023)
SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
by: Lim, Derek, et al.
Published: (2024)
by: Lim, Derek, et al.
Published: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Fairness Aware Reward Optimization
by: Choi, Ching Lam, et al.
Published: (2026)
by: Choi, Ching Lam, et al.
Published: (2026)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
CSRv2: Unlocking Ultra-Sparse Embeddings
by: Guo, Lixuan, et al.
Published: (2026)
by: Guo, Lixuan, et al.
Published: (2026)
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
by: Xu, Minghui, et al.
Published: (2026)
by: Xu, Minghui, et al.
Published: (2026)
Bridging Interpretability and Robustness Using LIME-Guided Model Refinement
by: Nayyem, Navid, et al.
Published: (2024)
by: Nayyem, Navid, et al.
Published: (2024)
Route Experts by Sequence, not by Token
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
The Exact Sample Complexity Gain from Invariances for Kernel Regression
by: Tahmasebi, Behrooz, et al.
Published: (2023)
by: Tahmasebi, Behrooz, et al.
Published: (2023)
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Model Merging in the Essential Subspace
by: Li, Longhua, et al.
Published: (2026)
by: Li, Longhua, et al.
Published: (2026)
Similar Items
-
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024) -
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025) -
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025) -
An Augmentation Overlap Theory of Contrastive Learning
by: Zhang, Qi, et al.
Published: (2025) -
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)