A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Fuente:
arXiv
Guardado en:
| Autores principales: | Tang, Yiming, Saini, Harshvardhan, Yao, Zhaoqian, Lin, Zheng, Liao, Yizhen, Cui, Jingyi, Wang, Yisen, Du, Mengnan, Liu, Dianbo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
por: Saini, Harshvardhan, et al.
Publicado: (2026)
por: Saini, Harshvardhan, et al.
Publicado: (2026)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
por: Saini, Harshvardhan, et al.
Publicado: (2026)
por: Saini, Harshvardhan, et al.
Publicado: (2026)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
por: Yao, Yifei, et al.
Publicado: (2025)
por: Yao, Yifei, et al.
Publicado: (2025)
An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
por: Cui, Jingyi, et al.
Publicado: (2025)
por: Cui, Jingyi, et al.
Publicado: (2025)
Biconvex Biclustering
por: Rosen, Sam, et al.
Publicado: (2026)
por: Rosen, Sam, et al.
Publicado: (2026)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
por: Tang, Yiming, et al.
Publicado: (2025)
por: Tang, Yiming, et al.
Publicado: (2025)
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
por: Cui, Jingyi, et al.
Publicado: (2025)
por: Cui, Jingyi, et al.
Publicado: (2025)
Disciplined Biconvex Programming
por: Zhu, Hao, et al.
Publicado: (2025)
por: Zhu, Hao, et al.
Publicado: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
por: Zhao, Haiyan, et al.
Publicado: (2025)
por: Zhao, Haiyan, et al.
Publicado: (2025)
CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis
por: Tang, Yiming, et al.
Publicado: (2025)
por: Tang, Yiming, et al.
Publicado: (2025)
LoRA Training in the NTK Regime has No Spurious Local Minima
por: Jang, Uijeong, et al.
Publicado: (2024)
por: Jang, Uijeong, et al.
Publicado: (2024)
Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
por: Liu, Xingtu
Publicado: (2021)
por: Liu, Xingtu
Publicado: (2021)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
por: Shu, Dong, et al.
Publicado: (2025)
por: Shu, Dong, et al.
Publicado: (2025)
Fast PINN Eigensolvers via Biconvex Reformulation
por: Banderwaar, Akshay Sai, et al.
Publicado: (2025)
por: Banderwaar, Akshay Sai, et al.
Publicado: (2025)
Nonnegative Low-rank Matrix Recovery Can Have Spurious Local Minima
por: Zhang, Richard Y.
Publicado: (2025)
por: Zhang, Richard Y.
Publicado: (2025)
A Non-Monotone Line-Search Method for Minimizing Functions with Spurious Local Minima
por: Aminifard, Zohreh, et al.
Publicado: (2025)
por: Aminifard, Zohreh, et al.
Publicado: (2025)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
por: Zhang, Qi, et al.
Publicado: (2024)
por: Zhang, Qi, et al.
Publicado: (2024)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
por: He, Zhengfu, et al.
Publicado: (2024)
por: He, Zhengfu, et al.
Publicado: (2024)
Unsupervised Concept Discovery Mitigates Spurious Correlations
por: Arefin, Md Rifat, et al.
Publicado: (2024)
por: Arefin, Md Rifat, et al.
Publicado: (2024)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
por: Tolooshams, Bahareh, et al.
Publicado: (2025)
por: Tolooshams, Bahareh, et al.
Publicado: (2025)
Auto-Calibration and Biconvex Compressive Sensing with Applications to Parallel MRI
por: Ni, Yuan, et al.
Publicado: (2024)
por: Ni, Yuan, et al.
Publicado: (2024)
What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based Agents
por: Xue, Zhaoqian, et al.
Publicado: (2024)
por: Xue, Zhaoqian, et al.
Publicado: (2024)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
por: He, Zirui, et al.
Publicado: (2025)
por: He, Zirui, et al.
Publicado: (2025)
Mechanistic Interpretability of ASR models using Sparse Autoencoders
por: Pluth, Dan, et al.
Publicado: (2026)
por: Pluth, Dan, et al.
Publicado: (2026)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
por: Zhang, Yi-Ge, et al.
Publicado: (2025)
por: Zhang, Yi-Ge, et al.
Publicado: (2025)
An Inclusive Theoretical Framework of Robust Supervised Contrastive Loss against Label Noise
por: Cui, Jingyi, et al.
Publicado: (2025)
por: Cui, Jingyi, et al.
Publicado: (2025)
DILA: Dictionary Label Attention for Mechanistic Interpretability in High-dimensional Multi-label Medical Coding Prediction
por: Wu, John, et al.
Publicado: (2024)
por: Wu, John, et al.
Publicado: (2024)
KnowThyself: An Agentic Assistant for LLM Interpretability
por: Prasai, Suraj, et al.
Publicado: (2025)
por: Prasai, Suraj, et al.
Publicado: (2025)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
por: Lee, Dongyeop, et al.
Publicado: (2025)
por: Lee, Dongyeop, et al.
Publicado: (2025)
Combating Spurious Correlations in Graph Interpretability via Self-Reflection
por: Cai, Kecheng, et al.
Publicado: (2026)
por: Cai, Kecheng, et al.
Publicado: (2026)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
por: Erdogan, Ege, et al.
Publicado: (2025)
por: Erdogan, Ege, et al.
Publicado: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
por: Tahimic, Kriz, et al.
Publicado: (2025)
por: Tahimic, Kriz, et al.
Publicado: (2025)
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
por: Lehn-Schiøler, William, et al.
Publicado: (2026)
por: Lehn-Schiøler, William, et al.
Publicado: (2026)
Winning Tickets as Axiomatic Minima: A Structural Interpretation
por: Bresciano, Claudio
Publicado: (2026)
por: Bresciano, Claudio
Publicado: (2026)
Fine-Grained Interpretation of Political Opinions in Large Language Models
por: Hu, Jingyu, et al.
Publicado: (2025)
por: Hu, Jingyu, et al.
Publicado: (2025)
Semi-Unified Sparse Dictionary Learning with Learnable Top-K LISTA and FISTA Encoders
por: Lin, Fengsheng, et al.
Publicado: (2025)
por: Lin, Fengsheng, et al.
Publicado: (2025)
YOLO-RD: Introducing Relevant and Compact Explicit Knowledge to YOLO by Retriever-Dictionary
por: Tsui, Hao-Tang, et al.
Publicado: (2024)
por: Tsui, Hao-Tang, et al.
Publicado: (2024)
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
por: Han, Jiaojiao, et al.
Publicado: (2025)
por: Han, Jiaojiao, et al.
Publicado: (2025)
Self‐Assembled Biconvex Microlens Array Using Chiral Ferroelectric Nematic Liquid Crystals
por: Kelum Perera, et al.
Publicado: (2024)
por: Kelum Perera, et al.
Publicado: (2024)
Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability
por: Shi, Yingdong, et al.
Publicado: (2025)
por: Shi, Yingdong, et al.
Publicado: (2025)
Ejemplares similares
-
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
por: Saini, Harshvardhan, et al.
Publicado: (2026) -
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
por: Saini, Harshvardhan, et al.
Publicado: (2026) -
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
por: Yao, Yifei, et al.
Publicado: (2025) -
An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
por: Cui, Jingyi, et al.
Publicado: (2025) -
Biconvex Biclustering
por: Rosen, Sam, et al.
Publicado: (2026)