Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
Fuente:
arXiv
Saved in:
| Main Authors: | Lim, Sungjun, Kim, Heedong, Lee, Andrew, Song, Kyungwoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sufficient Invariant Learning for Distribution Shift
by: Kim, Taero, et al.
Published: (2022)
by: Kim, Taero, et al.
Published: (2022)
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025)
by: Kang, Joonseong, et al.
Published: (2025)
Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
by: Choi, Youngjun, et al.
Published: (2025)
by: Choi, Youngjun, et al.
Published: (2025)
Uncertainty-driven Embedding Convolution
by: Lim, Sungjun, et al.
Published: (2025)
by: Lim, Sungjun, et al.
Published: (2025)
Faithful and Robust Local Interpretability for Textual Predictions
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
Semi-Supervised Preference Optimization with Limited Feedback
by: Lee, Seonggyun, et al.
Published: (2025)
by: Lee, Seonggyun, et al.
Published: (2025)
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing
by: Madsen, Andreas
Published: (2024)
by: Madsen, Andreas
Published: (2024)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
by: Han, Sungjun, et al.
Published: (2024)
by: Han, Sungjun, et al.
Published: (2024)
Transformer Explainer: Interactive Learning of Text-Generative Models
by: Cho, Aeree, et al.
Published: (2024)
by: Cho, Aeree, et al.
Published: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
Foundation Model's Embedded Representations May Detect Distribution Shift
by: Vargas, Max, et al.
Published: (2023)
by: Vargas, Max, et al.
Published: (2023)
When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift
by: Sheng, Zhecheng, et al.
Published: (2026)
by: Sheng, Zhecheng, et al.
Published: (2026)
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
by: Bao, Yuntai, et al.
Published: (2026)
by: Bao, Yuntai, et al.
Published: (2026)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
by: Kiruluta, Andrew
Published: (2025)
by: Kiruluta, Andrew
Published: (2025)
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026)
by: Le, Khoi, et al.
Published: (2026)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
by: Jegou, Simon, et al.
Published: (2026)
by: Jegou, Simon, et al.
Published: (2026)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
by: Jia, Jinghan, et al.
Published: (2026)
by: Jia, Jinghan, et al.
Published: (2026)
Shared Global and Local Geometry of Language Model Embeddings
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
by: Lee, Seongmin, et al.
Published: (2023)
by: Lee, Seongmin, et al.
Published: (2023)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
Locate&Edit: Energy-based Text Editing for Efficient, Flexible, and Faithful Controlled Text Generation
by: Son, Hye Ryung, et al.
Published: (2024)
by: Son, Hye Ryung, et al.
Published: (2024)
Causal Fine-Tuning under Latent Confounded Shift
by: Yu, Jialin, et al.
Published: (2024)
by: Yu, Jialin, et al.
Published: (2024)
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
by: Han, Songhao, et al.
Published: (2023)
by: Han, Songhao, et al.
Published: (2023)
Improving Rare Word Translation With Dictionaries and Attention Masking
by: Sible, Kenneth J., et al.
Published: (2024)
by: Sible, Kenneth J., et al.
Published: (2024)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
Mapping Faithful Reasoning in Language Models
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Multilingual Self-Taught Faithfulness Evaluators
by: Alfano, Carlo, et al.
Published: (2025)
by: Alfano, Carlo, et al.
Published: (2025)
MetaRM: Shifted Distributions Alignment via Meta-Learning
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
Interpretable inverse design of optical multilayer thin films based on extended neural adjoint and regression activation mapping
by: Kim, Sungjun, et al.
Published: (2025)
by: Kim, Sungjun, et al.
Published: (2025)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
by: Sim, Woo Seob, et al.
Published: (2026)
by: Sim, Woo Seob, et al.
Published: (2026)
A Bayesian Interpretation of Adaptive Low-Rank Adaptation
by: Chen, Haolin, et al.
Published: (2024)
by: Chen, Haolin, et al.
Published: (2024)
Similar Items
-
Sufficient Invariant Learning for Distribution Shift
by: Kim, Taero, et al.
Published: (2022) -
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025) -
Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
by: Choi, Youngjun, et al.
Published: (2025) -
Uncertainty-driven Embedding Convolution
by: Lim, Sungjun, et al.
Published: (2025) -
Faithful and Robust Local Interpretability for Textual Predictions
by: Lopardo, Gianluigi, et al.
Published: (2023)