LLMs Explain't: A Post-Mortem on Semantic Interpretability in Transformer Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdelhalim, Alhassan, Edinger, Janick, Laue, Sören, Regneri, Michaela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detecting Conceptual Abstraction in LLMs
von: Regneri, Michaela, et al.
Veröffentlicht: (2024)
von: Regneri, Michaela, et al.
Veröffentlicht: (2024)
Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2025)
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2025)
Automating Violence Detection and Categorization from Ancient Texts
von: Abdelhalim, Alhassan, et al.
Veröffentlicht: (2025)
von: Abdelhalim, Alhassan, et al.
Veröffentlicht: (2025)
Explaining Text Similarity in Transformer Models
von: Vasileiou, Alexandros, et al.
Veröffentlicht: (2024)
von: Vasileiou, Alexandros, et al.
Veröffentlicht: (2024)
Data Science with LLMs and Interpretable Models
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2024)
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2024)
Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2025)
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2025)
Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review
von: Gülsoylu, Emre, et al.
Veröffentlicht: (2025)
von: Gülsoylu, Emre, et al.
Veröffentlicht: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
von: Mihaila, George
Veröffentlicht: (2026)
von: Mihaila, George
Veröffentlicht: (2026)
Interpreting and Mitigating Unwanted Uncertainty in LLMs
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
A Structural Theory of Position Bias in Transformers
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2026)
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2026)
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
von: Zhang, Qingru, et al.
Veröffentlicht: (2023)
von: Zhang, Qingru, et al.
Veröffentlicht: (2023)
Correcting Suppressed Log-Probabilities in Language Models with Post-Transformer Adapters
von: Sanchez, Bryan
Veröffentlicht: (2026)
von: Sanchez, Bryan
Veröffentlicht: (2026)
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024)
von: Li, Jason
Veröffentlicht: (2024)
Efficient Line Search Method Based on Regression and Uncertainty Quantification
von: Laue, Sören, et al.
Veröffentlicht: (2024)
von: Laue, Sören, et al.
Veröffentlicht: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
LLMs for XAI: Future Directions for Explaining Explanations
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
A Short Survey of Human Mobility Prediction in Epidemic Modeling from Transformers to LLMs
von: Mayemba, Christian N., et al.
Veröffentlicht: (2024)
von: Mayemba, Christian N., et al.
Veröffentlicht: (2024)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025)
von: Jones, Erik, et al.
Veröffentlicht: (2025)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
von: Anagnostidis, Sotiris, et al.
Veröffentlicht: (2023)
von: Anagnostidis, Sotiris, et al.
Veröffentlicht: (2023)
Sparse Semantic Dimension as a Generalization Certificate for LLMs
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2026)
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2026)
Interpreting the Effects of Quantization on LLMs
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
von: Liu, Weiqi, et al.
Veröffentlicht: (2026)
von: Liu, Weiqi, et al.
Veröffentlicht: (2026)
Aligning Brain Activity with Advanced Transformer Models: Exploring the Role of Punctuation in Semantic Processing
von: Lamprou, Zenon, et al.
Veröffentlicht: (2025)
von: Lamprou, Zenon, et al.
Veröffentlicht: (2025)
Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
von: Nakkiran, Preetum, et al.
Veröffentlicht: (2025)
von: Nakkiran, Preetum, et al.
Veröffentlicht: (2025)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
von: Hemadri, Raghu Vamshi, et al.
Veröffentlicht: (2025)
von: Hemadri, Raghu Vamshi, et al.
Veröffentlicht: (2025)
The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
von: Kerce, J. Clayton, et al.
Veröffentlicht: (2026)
von: Kerce, J. Clayton, et al.
Veröffentlicht: (2026)
Fairness Definitions in Language Models Explained
von: Yin, Zhipeng, et al.
Veröffentlicht: (2024)
von: Yin, Zhipeng, et al.
Veröffentlicht: (2024)
DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs
von: Ganon, Ben, et al.
Veröffentlicht: (2024)
von: Ganon, Ben, et al.
Veröffentlicht: (2024)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
von: Gill, Waris, et al.
Veröffentlicht: (2025)
von: Gill, Waris, et al.
Veröffentlicht: (2025)
A General Framework for Producing Interpretable Semantic Text Embeddings
von: Sun, Yiqun, et al.
Veröffentlicht: (2024)
von: Sun, Yiqun, et al.
Veröffentlicht: (2024)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
von: Kiruluta, Andrew
Veröffentlicht: (2025)
von: Kiruluta, Andrew
Veröffentlicht: (2025)
JoPA:Explaining Large Language Model's Generation via Joint Prompt Attribution
von: Chang, Yurui, et al.
Veröffentlicht: (2024)
von: Chang, Yurui, et al.
Veröffentlicht: (2024)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
von: Gordon, Ofir, et al.
Veröffentlicht: (2026)
von: Gordon, Ofir, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Detecting Conceptual Abstraction in LLMs
von: Regneri, Michaela, et al.
Veröffentlicht: (2024) -
Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2025) -
Automating Violence Detection and Categorization from Ancient Texts
von: Abdelhalim, Alhassan, et al.
Veröffentlicht: (2025) -
Explaining Text Similarity in Transformer Models
von: Vasileiou, Alexandros, et al.
Veröffentlicht: (2024) -
Data Science with LLMs and Interpretable Models
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)