From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Mosbach, Marius, Gautam, Vagrant, Vergara-Browne, Tomás, Klakow, Dietrich, Geva, Mor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
The Hidden Space of Transformer Language Adapters
by: Alabi, Jesujoba O., et al.
Published: (2024)
by: Alabi, Jesujoba O., et al.
Published: (2024)
The Impact of Demonstrations on Multilingual In-Context Learning: A Multidimensional Analysis
by: Zhang, Miaoran, et al.
Published: (2024)
by: Zhang, Miaoran, et al.
Published: (2024)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
Teaching and Critiquing Conceptualization and Operationalization in NLP
by: Gautam, Vagrant
Published: (2025)
by: Gautam, Vagrant
Published: (2025)
Exploring the Effectiveness and Consistency of Task Selection in Intermediate-Task Transfer Learning
by: Lin, Pin-Jie, et al.
Published: (2024)
by: Lin, Pin-Jie, et al.
Published: (2024)
Aligned Probing: Relating Toxic Behavior and Model Internals
by: Waldis, Andreas, et al.
Published: (2025)
by: Waldis, Andreas, et al.
Published: (2025)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
by: García-de-Herreros, Paloma, et al.
Published: (2025)
by: García-de-Herreros, Paloma, et al.
Published: (2025)
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
by: Gottesman, Daniela, et al.
Published: (2025)
by: Gottesman, Daniela, et al.
Published: (2025)
Estimating Knowledge in Large Language Models Without Generating a Single Token
by: Gottesman, Daniela, et al.
Published: (2024)
by: Gottesman, Daniela, et al.
Published: (2024)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead
by: Alabi, Jesujoba O., et al.
Published: (2025)
by: Alabi, Jesujoba O., et al.
Published: (2025)
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
by: Mewes, Fabian, et al.
Published: (2026)
by: Mewes, Fabian, et al.
Published: (2026)
Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding
by: Blatt, Alexander, et al.
Published: (2024)
by: Blatt, Alexander, et al.
Published: (2024)
Tracr-Injection: Distilling Algorithms into Pre-trained Language Models
by: Vergara-Browne, Tomás, et al.
Published: (2025)
by: Vergara-Browne, Tomás, et al.
Published: (2025)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Your Multimodal Speech Model Says I Have a Face for Radio
by: Nachesa, Maya K., et al.
Published: (2026)
by: Nachesa, Maya K., et al.
Published: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
by: Ramati, Dana, et al.
Published: (2024)
by: Ramati, Dana, et al.
Published: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
by: Yona, Gal, et al.
Published: (2026)
by: Yona, Gal, et al.
Published: (2026)
IGC: Integrating a Gated Calculator into an LLM to Solve Arithmetic Tasks Reliably and Efficiently
by: Dietz, Florian, et al.
Published: (2025)
by: Dietz, Florian, et al.
Published: (2025)
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
by: Schuster, Jakob, et al.
Published: (2026)
by: Schuster, Jakob, et al.
Published: (2026)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
by: Rizvi-Martel, Michael, et al.
Published: (2026)
by: Rizvi-Martel, Michael, et al.
Published: (2026)
Not All Data Are Unlearned Equally
by: Krishnan, Aravind, et al.
Published: (2025)
by: Krishnan, Aravind, et al.
Published: (2025)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
by: Ivgi, Maor, et al.
Published: (2024)
by: Ivgi, Maor, et al.
Published: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
by: Barbi, Ohav, et al.
Published: (2025)
by: Barbi, Ohav, et al.
Published: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
by: Ahrac, Sagi, et al.
Published: (2026)
by: Ahrac, Sagi, et al.
Published: (2026)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026)
by: Gur-Arieh, Yoav, et al.
Published: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026)
by: Avrahamy, Asaf, et al.
Published: (2026)
Detecting (Un)answerability in Large Language Models with Linear Directions
by: Lavi, Maor Juliet, et al.
Published: (2025)
by: Lavi, Maor Juliet, et al.
Published: (2025)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
by: Shafran, Or, et al.
Published: (2026)
by: Shafran, Or, et al.
Published: (2026)
Similar Items
-
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024) -
The Hidden Space of Transformer Language Adapters
by: Alabi, Jesujoba O., et al.
Published: (2024) -
The Impact of Demonstrations on Multilingual In-Context Learning: A Multidimensional Analysis
by: Zhang, Miaoran, et al.
Published: (2024) -
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024) -
Teaching and Critiquing Conceptualization and Operationalization in NLP
by: Gautam, Vagrant
Published: (2025)