Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
Fuente:
arXiv
Saved in:
| Main Authors: | Lopardo, Gianluigi, Garreau, Damien |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
Faithful and Robust Local Interpretability for Textual Predictions
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
by: Lopardo, Gianluigi, et al.
Published: (2024)
by: Lopardo, Gianluigi, et al.
Published: (2024)
Rule2Text: Natural Language Explanation of Logical Rules in Knowledge Graphs
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
by: Guliani, Keerat, et al.
Published: (2026)
by: Guliani, Keerat, et al.
Published: (2026)
RuleR: Improving LLM Controllability by Rule-based Data Recycling
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
Student sentiment Analysis Using Classification With Feature Extraction Techniques
by: Tamrakar, Latika, et al.
Published: (2021)
by: Tamrakar, Latika, et al.
Published: (2021)
An Autoregressive Text-to-Graph Framework for Joint Entity and Relation Extraction
by: Zaratiana, Urchade, et al.
Published: (2024)
by: Zaratiana, Urchade, et al.
Published: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
by: Chawla, Krrish, et al.
Published: (2025)
by: Chawla, Krrish, et al.
Published: (2025)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
by: Miller, Moritz, et al.
Published: (2026)
by: Miller, Moritz, et al.
Published: (2026)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
by: Tuck, Bryan E., et al.
Published: (2025)
by: Tuck, Bryan E., et al.
Published: (2025)
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Interpretable Discriminative Text Representations via Agreement and Label Disentanglement
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
A General Framework for Producing Interpretable Semantic Text Embeddings
by: Sun, Yiqun, et al.
Published: (2024)
by: Sun, Yiqun, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
by: Li, Xiaomin, et al.
Published: (2024)
by: Li, Xiaomin, et al.
Published: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
by: Soo, Samuel, et al.
Published: (2025)
by: Soo, Samuel, et al.
Published: (2025)
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
by: Gao, Jiechao, et al.
Published: (2026)
by: Gao, Jiechao, et al.
Published: (2026)
Interpretable Predictability-Based AI Text Detection: A Replication Study
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
Rule by Rule: Learning with Confidence through Vocabulary Expansion
by: Nössig, Albert, et al.
Published: (2024)
by: Nössig, Albert, et al.
Published: (2024)
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
by: Butler, Landon, et al.
Published: (2025)
by: Butler, Landon, et al.
Published: (2025)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
by: Han, Sungmin, et al.
Published: (2025)
by: Han, Sungmin, et al.
Published: (2025)
Can Interpretation Predict Behavior on Unseen Data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
Chem-FINESE: Validating Fine-Grained Few-shot Entity Extraction through Text Reconstruction
by: Wang, Qingyun, et al.
Published: (2024)
by: Wang, Qingyun, et al.
Published: (2024)
Interpretable Recognition of Cognitive Distortions in Natural Language Texts
by: Kolonin, Anton, et al.
Published: (2025)
by: Kolonin, Anton, et al.
Published: (2025)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
by: Choi, Minsik, et al.
Published: (2025)
by: Choi, Minsik, et al.
Published: (2025)
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
by: Chen, Zheng, et al.
Published: (2025)
by: Chen, Zheng, et al.
Published: (2025)
GLEAMS: Bridging the Gap Between Local and Global Explanations
by: Visani, Giorgio, et al.
Published: (2024)
by: Visani, Giorgio, et al.
Published: (2024)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Health Insurance Coverage Rule Interpretation Corpus: Law, Policy, and Medical Guidance for Health Insurance Coverage Understanding
by: Gartner, Mike
Published: (2025)
by: Gartner, Mike
Published: (2025)
Detecting Suicidal Ideation in Text with Interpretable Deep Learning: A CNN-BiGRU with Attention Mechanism
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023)
by: Zhao, Zhixue, et al.
Published: (2023)
Similar Items
-
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022) -
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023) -
Faithful and Robust Local Interpretability for Textual Predictions
by: Lopardo, Gianluigi, et al.
Published: (2023) -
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
by: Lopardo, Gianluigi, et al.
Published: (2024) -
Rule2Text: Natural Language Explanation of Logical Rules in Knowledge Graphs
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)