Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
Fuente:
arXiv
Saved in:
| Main Authors: | Oldfield, James, Im, Shawn, Li, Sharon, Nicolaou, Mihalis A., Patras, Ioannis, Chrysos, Grigorios G |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
by: Oldfield, James, et al.
Published: (2024)
by: Oldfield, James, et al.
Published: (2024)
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
by: Koromilas, Panagiotis, et al.
Published: (2026)
by: Koromilas, Panagiotis, et al.
Published: (2026)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)
by: Afzal, Arshia, et al.
Published: (2024)
fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery
by: Demou, Andreas D., et al.
Published: (2026)
by: Demou, Andreas D., et al.
Published: (2026)
Motor Imagery Decoding Using Ensemble Curriculum Learning and Collaborative Training
by: Zoumpourlis, Georgios, et al.
Published: (2022)
by: Zoumpourlis, Georgios, et al.
Published: (2022)
Robust NAS under adversarial training: benchmark, theory, and beyond
by: Wu, Yongtao, et al.
Published: (2024)
by: Wu, Yongtao, et al.
Published: (2024)
Understanding the Learning Dynamics of Alignment with Human Feedback
by: Im, Shawn, et al.
Published: (2024)
by: Im, Shawn, et al.
Published: (2024)
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
by: Candogan, Leyla Naz, et al.
Published: (2025)
by: Candogan, Leyla Naz, et al.
Published: (2025)
Efficient local linearity regularization to overcome catastrophic overfitting
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
Why DDIM Hallucinates More Than DDPM: A Theoretical Analysis of Reverse Dynamics
by: Ashiq, Muhammad H., et al.
Published: (2026)
by: Ashiq, Muhammad H., et al.
Published: (2026)
Neural Networks Decoded: Targeted and Robust Analysis of Neural Network Decisions via Causal Explanations and Reasoning
by: Diallo, Alec F., et al.
Published: (2024)
by: Diallo, Alec F., et al.
Published: (2024)
Faithful Interpretation for Graph Neural Networks
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Multimodal Machine Learning in Mental Health: A Survey of Data, Algorithms, and Challenges
by: Sahili, Zahraa Al, et al.
Published: (2024)
by: Sahili, Zahraa Al, et al.
Published: (2024)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
by: Sahili, Zahraa Al, et al.
Published: (2024)
by: Sahili, Zahraa Al, et al.
Published: (2024)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained Encoders
by: Lei, Zhuxin, et al.
Published: (2026)
by: Lei, Zhuxin, et al.
Published: (2026)
Visual Instruction Bottleneck Tuning
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
Towards Metric-Faithful Neural Graph Matching
by: Shivottam, Jyotirmaya, et al.
Published: (2026)
by: Shivottam, Jyotirmaya, et al.
Published: (2026)
Mixture of Layers with Hybrid Attention
by: Ternovtsii, Ivan, et al.
Published: (2026)
by: Ternovtsii, Ivan, et al.
Published: (2026)
EUGens: Efficient, Unified, and General Dense Layers
by: Kim, Sang Min, et al.
Published: (2026)
by: Kim, Sang Min, et al.
Published: (2026)
Towards Faithful Explanations: Boosting Rationalization with Shortcuts Discovery
by: Yue, Linan, et al.
Published: (2024)
by: Yue, Linan, et al.
Published: (2024)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
Going beyond Compositions, DDPMs Can Produce Zero-Shot Interpolations
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
by: Huang, Xin, et al.
Published: (2026)
by: Huang, Xin, et al.
Published: (2026)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
by: Singh, Abhishek Kumar, et al.
Published: (2024)
by: Singh, Abhishek Kumar, et al.
Published: (2024)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
by: Jeon, Wonseok, et al.
Published: (2024)
by: Jeon, Wonseok, et al.
Published: (2024)
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
by: Novikov, Ivan
Published: (2025)
by: Novikov, Ivan
Published: (2025)
Interpretable Diffusion via Information Decomposition
by: Kong, Xianghao, et al.
Published: (2023)
by: Kong, Xianghao, et al.
Published: (2023)
Similar Items
-
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
by: Oldfield, James, et al.
Published: (2024) -
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025) -
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
by: Cheng, Yixin, et al.
Published: (2024) -
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
by: Koromilas, Panagiotis, et al.
Published: (2026) -
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)