ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
Fuente:
arXiv
Salvato in:
| Autori principali: | Poché, Antonin, Jacovi, Alon, Picard, Agustin Martin, Boutin, Victor, Jourdan, Fanny |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
di: Bernas, Raphael, et al.
Pubblicazione: (2026)
di: Bernas, Raphael, et al.
Pubblicazione: (2026)
TaCo: Targeted Concept Erasure Prevents Non-Linear Classifiers From Detecting Protected Attributes
di: Jourdan, Fanny, et al.
Pubblicazione: (2023)
di: Jourdan, Fanny, et al.
Pubblicazione: (2023)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
di: Jourdan, Fanny
Pubblicazione: (2024)
di: Jourdan, Fanny
Pubblicazione: (2024)
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning
di: Fouilhé, Guilhem, et al.
Pubblicazione: (2026)
di: Fouilhé, Guilhem, et al.
Pubblicazione: (2026)
Interpreto: An Explainability Library for Transformers
di: Poché, Antonin, et al.
Pubblicazione: (2025)
di: Poché, Antonin, et al.
Pubblicazione: (2025)
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
di: Hong, Pingjun, et al.
Pubblicazione: (2026)
di: Hong, Pingjun, et al.
Pubblicazione: (2026)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
di: Jourdan, Fanny, et al.
Pubblicazione: (2025)
di: Jourdan, Fanny, et al.
Pubblicazione: (2025)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
di: Goldman, Omer, et al.
Pubblicazione: (2024)
di: Goldman, Omer, et al.
Pubblicazione: (2024)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
di: Mills, Edmund, et al.
Pubblicazione: (2023)
di: Mills, Edmund, et al.
Pubblicazione: (2023)
Magnetism of stable structures of small binary ConSim (n + m < 4) clusters
di: E.M Sosa-Hernández
Pubblicazione: (2008)
di: E.M Sosa-Hernández
Pubblicazione: (2008)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
di: Dhimoïla, Grégoire, et al.
Pubblicazione: (2026)
di: Dhimoïla, Grégoire, et al.
Pubblicazione: (2026)
TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
di: Caciularu, Avi, et al.
Pubblicazione: (2024)
di: Caciularu, Avi, et al.
Pubblicazione: (2024)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
CoverBench: A Challenging Benchmark for Complex Claim Verification
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
Bisimilarity and Simulatability of Processes Parameterized by Join Interactions
di: Grabmayer, Clemens, et al.
Pubblicazione: (2025)
di: Grabmayer, Clemens, et al.
Pubblicazione: (2025)
ConQuer: A Framework for Concept-Based Quiz Generation
di: Fu, Yicheng, et al.
Pubblicazione: (2025)
di: Fu, Yicheng, et al.
Pubblicazione: (2025)
Latent Concept-based Explanation of NLP Models
di: Yu, Xuemin, et al.
Pubblicazione: (2024)
di: Yu, Xuemin, et al.
Pubblicazione: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
di: Alon, Bar, et al.
Pubblicazione: (2026)
di: Alon, Bar, et al.
Pubblicazione: (2026)
ConTrans: Weak-to-Strong Alignment Engineering via Concept Transplantation
di: Dong, Weilong, et al.
Pubblicazione: (2024)
di: Dong, Weilong, et al.
Pubblicazione: (2024)
ConExion: Concept Extraction with Large Language Models
di: Norouzi, Ebrahim, et al.
Pubblicazione: (2025)
di: Norouzi, Ebrahim, et al.
Pubblicazione: (2025)
DoubleDipper: Improving Long-Context LLMs via Context Recycling
di: Cattan, Arie, et al.
Pubblicazione: (2024)
di: Cattan, Arie, et al.
Pubblicazione: (2024)
Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery
di: Sun, Yifan, et al.
Pubblicazione: (2025)
di: Sun, Yifan, et al.
Pubblicazione: (2025)
Explanation sensitivity to the randomness of large language models: the case of journalistic text classification
di: Bogaert, Jeremie, et al.
Pubblicazione: (2024)
di: Bogaert, Jeremie, et al.
Pubblicazione: (2024)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
di: Cattan, Arie, et al.
Pubblicazione: (2025)
di: Cattan, Arie, et al.
Pubblicazione: (2025)
ConStat: Performance-Based Contamination Detection in Large Language Models
di: Dekoninck, Jasper, et al.
Pubblicazione: (2024)
di: Dekoninck, Jasper, et al.
Pubblicazione: (2024)
FGR-ColBERT: Identifying Fine-Grained Relevance Tokens During Retrieval
di: Jarolím, Antonín, et al.
Pubblicazione: (2026)
di: Jarolím, Antonín, et al.
Pubblicazione: (2026)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
di: Toker, Gilat, et al.
Pubblicazione: (2026)
di: Toker, Gilat, et al.
Pubblicazione: (2026)
Can LLMs extract human-like fine-grained evidence for evidence-based fact-checking?
di: Jarolím, Antonín, et al.
Pubblicazione: (2025)
di: Jarolím, Antonín, et al.
Pubblicazione: (2025)
A Straightforward Pipeline for Targeted Entailment and Contradiction Detection
di: Sulc, Antonin
Pubblicazione: (2025)
di: Sulc, Antonin
Pubblicazione: (2025)
Towards a Framework for Evaluating Explanations in Automated Fact Verification
di: Kotonya, Neema, et al.
Pubblicazione: (2024)
di: Kotonya, Neema, et al.
Pubblicazione: (2024)
Zonkey: A Hierarchical Diffusion Language Model with Differentiable Tokenization and Probabilistic Attention
di: Rozental, Alon
Pubblicazione: (2026)
di: Rozental, Alon
Pubblicazione: (2026)
Estimation of Concept Explanations Should be Uncertainty Aware
di: Piratla, Vihari, et al.
Pubblicazione: (2023)
di: Piratla, Vihari, et al.
Pubblicazione: (2023)
Event2Vec: A Geometric Approach to Learning Composable Representations of Event Sequences
di: Sulc, Antonin
Pubblicazione: (2025)
di: Sulc, Antonin
Pubblicazione: (2025)
Effective Explanations Support Planning Under Uncertainty
di: Zhou, Hanqi, et al.
Pubblicazione: (2026)
di: Zhou, Hanqi, et al.
Pubblicazione: (2026)
AutoPCR: Automated Phenotype Concept Recognition by Prompting
di: Tao, Yicheng, et al.
Pubblicazione: (2025)
di: Tao, Yicheng, et al.
Pubblicazione: (2025)
Automate Knowledge Concept Tagging on Math Questions with LLMs
di: Li, Hang, et al.
Pubblicazione: (2024)
di: Li, Hang, et al.
Pubblicazione: (2024)
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
di: Pala, Tej Deep, et al.
Pubblicazione: (2024)
di: Pala, Tej Deep, et al.
Pubblicazione: (2024)
LLM-Augmented Changepoint Detection: A Framework for Ensemble Detection and Automated Explanation
di: Lukassen, Fabian, et al.
Pubblicazione: (2026)
di: Lukassen, Fabian, et al.
Pubblicazione: (2026)
Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
di: Makaiova, Lucia, et al.
Pubblicazione: (2025)
di: Makaiova, Lucia, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
di: Bernas, Raphael, et al.
Pubblicazione: (2026) -
TaCo: Targeted Concept Erasure Prevents Non-Linear Classifiers From Detecting Protected Attributes
di: Jourdan, Fanny, et al.
Pubblicazione: (2023) -
Counterfactual Simulatability of LLM Explanations for Generation Tasks
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025) -
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
di: Jourdan, Fanny
Pubblicazione: (2024) -
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning
di: Fouilhé, Guilhem, et al.
Pubblicazione: (2026)