Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
Fuente:
arXiv
Guardado en:
| Autores principales: | Ayonrinde, Kola, Jaburi, Louis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
por: Ayonrinde, Kola, et al.
Publicado: (2025)
por: Ayonrinde, Kola, et al.
Publicado: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
por: Zeng, Qiuhai, et al.
Publicado: (2025)
por: Zeng, Qiuhai, et al.
Publicado: (2025)
LLMs for XAI: Future Directions for Explaining Explanations
por: Zytek, Alexandra, et al.
Publicado: (2024)
por: Zytek, Alexandra, et al.
Publicado: (2024)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
por: Lee, Seongmin, et al.
Publicado: (2023)
por: Lee, Seongmin, et al.
Publicado: (2023)
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
por: Wu, Yang, et al.
Publicado: (2026)
por: Wu, Yang, et al.
Publicado: (2026)
A Hopfieldian View-based Interpretation for Chain-of-Thought Reasoning
por: Hu, Lijie, et al.
Publicado: (2024)
por: Hu, Lijie, et al.
Publicado: (2024)
Properties and Challenges of LLM-Generated Explanations
por: Kunz, Jenny, et al.
Publicado: (2024)
por: Kunz, Jenny, et al.
Publicado: (2024)
Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications
por: Sakib, Abu Noman Md, et al.
Publicado: (2026)
por: Sakib, Abu Noman Md, et al.
Publicado: (2026)
Diagrammatization and Abduction to Improve AI Interpretability With Domain-Aligned Explanations for Medical Diagnosis
por: Lim, Brian Y., et al.
Publicado: (2023)
por: Lim, Brian Y., et al.
Publicado: (2023)
Heterogeneous Value Alignment Evaluation for Large Language Models
por: Zhang, Zhaowei, et al.
Publicado: (2023)
por: Zhang, Zhaowei, et al.
Publicado: (2023)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
por: Fu, Xiao, et al.
Publicado: (2025)
por: Fu, Xiao, et al.
Publicado: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
por: Kaur, Navreet, et al.
Publicado: (2023)
por: Kaur, Navreet, et al.
Publicado: (2023)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
por: Lee, Dong Won, et al.
Publicado: (2024)
por: Lee, Dong Won, et al.
Publicado: (2024)
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
por: Sun, Yuxuan, et al.
Publicado: (2025)
por: Sun, Yuxuan, et al.
Publicado: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
por: Cook, Jonathan, et al.
Publicado: (2024)
por: Cook, Jonathan, et al.
Publicado: (2024)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
por: Kahng, Minsuk, et al.
Publicado: (2024)
por: Kahng, Minsuk, et al.
Publicado: (2024)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
por: Ben-Zion, Ziv, et al.
Publicado: (2025)
por: Ben-Zion, Ziv, et al.
Publicado: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
por: Park, Jung In, et al.
Publicado: (2024)
por: Park, Jung In, et al.
Publicado: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
por: Baidya, Avinash, et al.
Publicado: (2025)
por: Baidya, Avinash, et al.
Publicado: (2025)
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
por: Xu, Ziwen, et al.
Publicado: (2026)
por: Xu, Ziwen, et al.
Publicado: (2026)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
por: Kim, Jane Paik
Publicado: (2026)
por: Kim, Jane Paik
Publicado: (2026)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
por: Soni, Nikita, et al.
Publicado: (2025)
por: Soni, Nikita, et al.
Publicado: (2025)
Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI
por: Saha, Agnik, et al.
Publicado: (2025)
por: Saha, Agnik, et al.
Publicado: (2025)
PRECISE Framework: GPT-based Text For Improved Readability, Reliability, and Understandability of Radiology Reports For Patient-Centered Care
por: Tripathi, Satvik, et al.
Publicado: (2024)
por: Tripathi, Satvik, et al.
Publicado: (2024)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
por: Sivaprasad, Adarsa, et al.
Publicado: (2023)
por: Sivaprasad, Adarsa, et al.
Publicado: (2023)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
por: Rajagopal, Shreya, et al.
Publicado: (2024)
por: Rajagopal, Shreya, et al.
Publicado: (2024)
UniAutoML: A Human-Centered Framework for Unified Discriminative and Generative AutoML with Large Language Models
por: Guo, Jiayi, et al.
Publicado: (2024)
por: Guo, Jiayi, et al.
Publicado: (2024)
Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding
por: Baer, Gregor, et al.
Publicado: (2026)
por: Baer, Gregor, et al.
Publicado: (2026)
TalkWithMachines: Enhancing Human-Robot Interaction for Interpretable Industrial Robotics Through Large/Vision Language Models
por: Abbas, Ammar N., et al.
Publicado: (2024)
por: Abbas, Ammar N., et al.
Publicado: (2024)
MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems
por: Chari, Shruthi, et al.
Publicado: (2025)
por: Chari, Shruthi, et al.
Publicado: (2025)
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
por: Feher, Darius, et al.
Publicado: (2024)
por: Feher, Darius, et al.
Publicado: (2024)
Predicting Satisfaction of Counterfactual Explanations from Human Ratings of Explanatory Qualities
por: Domnich, Marharyta, et al.
Publicado: (2025)
por: Domnich, Marharyta, et al.
Publicado: (2025)
EXMOS: Explanatory Model Steering Through Multifaceted Explanations and Data Configurations
por: Bhattacharya, Aditya, et al.
Publicado: (2024)
por: Bhattacharya, Aditya, et al.
Publicado: (2024)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
por: Chaleshtori, Fateme Hashemi, et al.
Publicado: (2024)
por: Chaleshtori, Fateme Hashemi, et al.
Publicado: (2024)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
por: Sung, Yoo Yeon, et al.
Publicado: (2025)
por: Sung, Yoo Yeon, et al.
Publicado: (2025)
User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
por: Carvallo, Andrés, et al.
Publicado: (2025)
por: Carvallo, Andrés, et al.
Publicado: (2025)
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
por: Ou, Yixin, et al.
Publicado: (2025)
por: Ou, Yixin, et al.
Publicado: (2025)
From Feature Importance to Natural Language Explanations Using LLMs with RAG
por: Tekkesinoglu, Sule, et al.
Publicado: (2024)
por: Tekkesinoglu, Sule, et al.
Publicado: (2024)
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
por: Karamolegkou, Antonia, et al.
Publicado: (2025)
por: Karamolegkou, Antonia, et al.
Publicado: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
por: Sturgeon, Benjamin, et al.
Publicado: (2025)
por: Sturgeon, Benjamin, et al.
Publicado: (2025)
Ejemplares similares
-
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
por: Ayonrinde, Kola, et al.
Publicado: (2025) -
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
por: Zeng, Qiuhai, et al.
Publicado: (2025) -
LLMs for XAI: Future Directions for Explaining Explanations
por: Zytek, Alexandra, et al.
Publicado: (2024) -
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
por: Lee, Seongmin, et al.
Publicado: (2023) -
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
por: Wu, Yang, et al.
Publicado: (2026)