ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Butler, Landon, Agarwal, Abhineet, Kang, Justin Singh, Erginbas, Yigit Efe, Yu, Bin, Ramchandran, Kannan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Learning to Understand: Identifying Interactions via the Möbius Transform
by: Kang, Justin S., et al.
Published: (2024)
by: Kang, Justin S., et al.
Published: (2024)
Adaptive Sparse Möbius Transforms for Learning Polynomials
by: Erginbas, Yigit Efe, et al.
Published: (2026)
by: Erginbas, Yigit Efe, et al.
Published: (2026)
Online Assortment and Price Optimization Under Contextual Choice Models
by: Erginbas, Yigit Efe, et al.
Published: (2025)
by: Erginbas, Yigit Efe, et al.
Published: (2025)
An Odd Estimator for Shapley Values
by: Fumagalli, Fabian, et al.
Published: (2026)
by: Fumagalli, Fabian, et al.
Published: (2026)
SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future Queries
by: Tsui, Darin, et al.
Published: (2024)
by: Tsui, Darin, et al.
Published: (2024)
The Fair Value of Data Under Heterogeneous Privacy Constraints in Federated Learning
by: Kang, Justin, et al.
Published: (2023)
by: Kang, Justin, et al.
Published: (2023)
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Synthetic Combinations: A Causal Inference Framework for Combinatorial Interventions
by: Agarwal, Abhineet, et al.
Published: (2023)
by: Agarwal, Abhineet, et al.
Published: (2023)
Multi-Armed Bandits with Network Interference
by: Agarwal, Abhineet, et al.
Published: (2024)
by: Agarwal, Abhineet, et al.
Published: (2024)
Local MDI+: Local Feature Importances for Tree-Based Models
by: Liang, Zhongyuan, et al.
Published: (2025)
by: Liang, Zhongyuan, et al.
Published: (2025)
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Adaptive Test-Time Intervention for Concept Bottleneck Models
by: Shen, Matthew, et al.
Published: (2025)
by: Shen, Matthew, et al.
Published: (2025)
Integrating Random Forests and Generalized Linear Models for Improved Accuracy and Interpretability
by: Agarwal, Abhineet, et al.
Published: (2023)
by: Agarwal, Abhineet, et al.
Published: (2023)
Tokenizing Semantic Segmentation with Run Length Encoding
by: Singh, Abhineet, et al.
Published: (2026)
by: Singh, Abhineet, et al.
Published: (2026)
Statistical Proof of Execution (SPEX)
by: Dallachiesa, Michele, et al.
Published: (2025)
by: Dallachiesa, Michele, et al.
Published: (2025)
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
by: Liu, Junhao, et al.
Published: (2025)
by: Liu, Junhao, et al.
Published: (2025)
PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework
by: Agarwal, Abhineet, et al.
Published: (2025)
by: Agarwal, Abhineet, et al.
Published: (2025)
SAGE: A Realistic Benchmark for Semantic Understanding
by: Goel, Samarth, et al.
Published: (2025)
by: Goel, Samarth, et al.
Published: (2025)
Quantifying Positional Biases in Text Embedding Models
by: Lee, Reagan J., et al.
Published: (2024)
by: Lee, Reagan J., et al.
Published: (2024)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
by: Sun, Liwen, et al.
Published: (2024)
by: Sun, Liwen, et al.
Published: (2024)
Mapping plasma properties of Cassiopeia A with XRISM/Resolve: a Bayesian analysis via UltraSPEX
by: Agarwal, Manan, et al.
Published: (2026)
by: Agarwal, Manan, et al.
Published: (2026)
Improving Token-based Object Detection with Video
by: Singh, Abhineet, et al.
Published: (2025)
by: Singh, Abhineet, et al.
Published: (2025)
AgentSPEX: An Agent SPecification and EXecution Language
by: Wang, Pengcheng, et al.
Published: (2026)
by: Wang, Pengcheng, et al.
Published: (2026)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023)
by: Rajaraman, Nived, et al.
Published: (2023)
Looped Transformers for Length Generalization
by: Fan, Ying, et al.
Published: (2024)
by: Fan, Ying, et al.
Published: (2024)
Mixture-of-Channels: Exploiting Sparse FFNs for Efficient LLMs Pre-Training and Inference
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Interpretable and Testable Vision Features via Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
Double Blind Imaging with Generative Modeling
by: Levac, Brett, et al.
Published: (2025)
by: Levac, Brett, et al.
Published: (2025)
Interpret and Control Dense Retrieval with Sparse Latent Features
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Towards Anytime-Valid Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2026)
by: Huang, Baihe, et al.
Published: (2026)
Analog Alchemy: Neural Computation with In-Memory Inference, Learning and Routing
by: Demirag, Yigit
Published: (2024)
by: Demirag, Yigit
Published: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
by: Tahimic, Kriz, et al.
Published: (2025)
by: Tahimic, Kriz, et al.
Published: (2025)
A Mathematical Model to Capture Urbanization Trajectory Induced by Economic Inequality
by: Pandey, Neeraj, et al.
Published: (2025)
by: Pandey, Neeraj, et al.
Published: (2025)
EmbedLLM: Learning Compact Representations of Large Language Models
by: Zhuang, Richard, et al.
Published: (2024)
by: Zhuang, Richard, et al.
Published: (2024)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
by: Fei, Yu, et al.
Published: (2024)
by: Fei, Yu, et al.
Published: (2024)
ProxyAttn: Guided Sparse Attention via Representative Heads
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Similar Items
-
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025) -
Learning to Understand: Identifying Interactions via the Möbius Transform
by: Kang, Justin S., et al.
Published: (2024) -
Adaptive Sparse Möbius Transforms for Learning Polynomials
by: Erginbas, Yigit Efe, et al.
Published: (2026) -
Online Assortment and Price Optimization Under Contextual Choice Models
by: Erginbas, Yigit Efe, et al.
Published: (2025) -
An Odd Estimator for Shapley Values
by: Fumagalli, Fabian, et al.
Published: (2026)