Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Ziqian, Muhamed, Aashiq, Diab, Mona T., Smith, Virginia, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
CoRAG: Collaborative Retrieval-Augmented Generation
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026)
by: Wedgwood, James, et al.
Published: (2026)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
by: Muhamed, Aashiq
Published: (2025)
by: Muhamed, Aashiq
Published: (2025)
Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Hodoscope: Unsupervised Monitoring for AI Misbehaviors
by: Zhong, Ziqian, et al.
Published: (2026)
by: Zhong, Ziqian, et al.
Published: (2026)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
by: Khullar, Dipika, et al.
Published: (2026)
by: Khullar, Dipika, et al.
Published: (2026)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
Reasoning Traces Shape Outputs but Models Won't Say So
by: Hao, Yijie, et al.
Published: (2026)
by: Hao, Yijie, et al.
Published: (2026)
Reasoning as an Adaptive Defense for Safety
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
by: Dadfar, Zachary Pedram
Published: (2026)
by: Dadfar, Zachary Pedram
Published: (2026)
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024)
by: Tafreshi, Shabnam, et al.
Published: (2024)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Mode-Conditioning Unlocks Superior Test-Time Scaling
by: Wu, Chen Henry, et al.
Published: (2025)
by: Wu, Chen Henry, et al.
Published: (2025)
The World Won't Stay Still: Programmable Evolution for Agent Benchmarks
by: Li, Guangrui, et al.
Published: (2026)
by: Li, Guangrui, et al.
Published: (2026)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
by: Kim, Eungyeup, et al.
Published: (2023)
by: Kim, Eungyeup, et al.
Published: (2023)
S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
by: Chavan, Arnav, et al.
Published: (2026)
by: Chavan, Arnav, et al.
Published: (2026)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
by: Hao, Chenjie, et al.
Published: (2026)
by: Hao, Chenjie, et al.
Published: (2026)
Explaining the Explainer: Understanding the Inner Workings of Transformer-based Symbolic Regression Models
by: van Breda, Arco, et al.
Published: (2026)
by: van Breda, Arco, et al.
Published: (2026)
Explaining the Unexplained: Revealing Hidden Correlations for Better Interpretability
by: Jiang, Wen-Dong, et al.
Published: (2024)
by: Jiang, Wen-Dong, et al.
Published: (2024)
Algorithmic Capabilities of Random Transformers
by: Zhong, Ziqian, et al.
Published: (2024)
by: Zhong, Ziqian, et al.
Published: (2024)
Explaining Machine Learning Predictive Models through Conditional Expectation Methods
by: Ruiz-España, Silvia, et al.
Published: (2026)
by: Ruiz-España, Silvia, et al.
Published: (2026)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
by: Deng, Yihe, et al.
Published: (2023)
by: Deng, Yihe, et al.
Published: (2023)
Reference-based Metrics Disprove Themselves in Question Generation
by: Nguyen, Bang, et al.
Published: (2024)
by: Nguyen, Bang, et al.
Published: (2024)
Position: Do Not Explain Vision Models Without Context
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
Similar Items
-
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024) -
CoRAG: Collaborative Retrieval-Augmented Generation
by: Muhamed, Aashiq, et al.
Published: (2025) -
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026) -
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025) -
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)