Evaluating Model Explanations without Ground Truth
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rawal, Kaivalya, Fu, Zihao, Delaney, Eoin, Russell, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
von: Rawal, Kaivalya, et al.
Veröffentlicht: (2026)
von: Rawal, Kaivalya, et al.
Veröffentlicht: (2026)
Choosing DAG Models Using Markov and Minimal Edge Count in the Absence of Ground Truth
von: Ramsey, Joseph D., et al.
Veröffentlicht: (2024)
von: Ramsey, Joseph D., et al.
Veröffentlicht: (2024)
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
von: Ribeiro, José, et al.
Veröffentlicht: (2022)
von: Ribeiro, José, et al.
Veröffentlicht: (2022)
How Reliable and Stable are Explanations of XAI Methods?
von: Ribeiro, José, et al.
Veröffentlicht: (2024)
von: Ribeiro, José, et al.
Veröffentlicht: (2024)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
von: Gururaj, Shreyas, et al.
Veröffentlicht: (2025)
von: Gururaj, Shreyas, et al.
Veröffentlicht: (2025)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites
von: Zhang, Qinshi, et al.
Veröffentlicht: (2026)
von: Zhang, Qinshi, et al.
Veröffentlicht: (2026)
Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
von: Li, Andrew C., et al.
Veröffentlicht: (2025)
von: Li, Andrew C., et al.
Veröffentlicht: (2025)
FORCE: Feature-Oriented Representation with Clustering and Explanation
von: Mukherjee, Rishav, et al.
Veröffentlicht: (2025)
von: Mukherjee, Rishav, et al.
Veröffentlicht: (2025)
Algebraic Evaluation Theorems
von: Corrada-Emmanuel, Andrés
Veröffentlicht: (2024)
von: Corrada-Emmanuel, Andrés
Veröffentlicht: (2024)
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
von: Carrow, Stephen, et al.
Veröffentlicht: (2024)
von: Carrow, Stephen, et al.
Veröffentlicht: (2024)
Evaluation of post-hoc interpretability methods in time-series classification
von: Turbé, Hugues, et al.
Veröffentlicht: (2022)
von: Turbé, Hugues, et al.
Veröffentlicht: (2022)
Evaluating the effectiveness of predicting covariates in LSTM Networks for Time Series Forecasting
von: Davies, Gareth
Veröffentlicht: (2024)
von: Davies, Gareth
Veröffentlicht: (2024)
Leveraging Personalized PageRank and Higher-Order Topological Structures for Heterophily Mitigation in Graph Neural Networks
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2025)
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2025)
Incentives for Responsiveness, Instrumental Control and Impact
von: Carey, Ryan, et al.
Veröffentlicht: (2020)
von: Carey, Ryan, et al.
Veröffentlicht: (2020)
Large Language Models as Attribution Regularizers for Efficient Model Training
von: Vukadin, Davor, et al.
Veröffentlicht: (2025)
von: Vukadin, Davor, et al.
Veröffentlicht: (2025)
Approximate Domain Unlearning for Vision-Language Models
von: Kawamura, Kodai, et al.
Veröffentlicht: (2025)
von: Kawamura, Kodai, et al.
Veröffentlicht: (2025)
Grokking Beyond the Euclidean Norm of Model Parameters
von: Notsawo, Pascal Jr Tikeng, et al.
Veröffentlicht: (2025)
von: Notsawo, Pascal Jr Tikeng, et al.
Veröffentlicht: (2025)
Scaling Offline RL via Efficient and Expressive Shortcut Models
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
Architectural Proprioception in State Space Models: Thermodynamic Training Induces Anticipatory Halt Detection
von: Noon, Jay
Veröffentlicht: (2026)
von: Noon, Jay
Veröffentlicht: (2026)
A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences
von: Siow, Bryan Y.
Veröffentlicht: (2025)
von: Siow, Bryan Y.
Veröffentlicht: (2025)
First-Mover Bias in Gradient Boosting Explanations: Mechanism, Detection, and Resolution
von: Caraker, Drake, et al.
Veröffentlicht: (2026)
von: Caraker, Drake, et al.
Veröffentlicht: (2026)
Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction
von: Kohlberger, Björn Roman
Veröffentlicht: (2026)
von: Kohlberger, Björn Roman
Veröffentlicht: (2026)
Model Fusion via Retrofitting
von: Luenam, Phoomraphee, et al.
Veröffentlicht: (2025)
von: Luenam, Phoomraphee, et al.
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
von: Anderson, Samuel Cyrenius
Veröffentlicht: (2026)
von: Anderson, Samuel Cyrenius
Veröffentlicht: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
von: Deshmukh, Pratik, et al.
Veröffentlicht: (2026)
von: Deshmukh, Pratik, et al.
Veröffentlicht: (2026)
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
HEHRGNN: A Unified Embedding Model for Knowledge Graphs with Hyperedges and Hyper-Relational Edges
von: Rajagopalamenon, Rajesh, et al.
Veröffentlicht: (2026)
von: Rajagopalamenon, Rajesh, et al.
Veröffentlicht: (2026)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
von: Rath, Plawan Kumar, et al.
Veröffentlicht: (2026)
von: Rath, Plawan Kumar, et al.
Veröffentlicht: (2026)
Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
von: Alsheikh, Ahmad, et al.
Veröffentlicht: (2025)
von: Alsheikh, Ahmad, et al.
Veröffentlicht: (2025)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
von: Zhang, Xinyu
Veröffentlicht: (2026)
von: Zhang, Xinyu
Veröffentlicht: (2026)
TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive Correction
von: Liu, Weijie, et al.
Veröffentlicht: (2025)
von: Liu, Weijie, et al.
Veröffentlicht: (2025)
DataRater: Meta-Learned Dataset Curation
von: Calian, Dan A., et al.
Veröffentlicht: (2025)
von: Calian, Dan A., et al.
Veröffentlicht: (2025)
The Lattice Geometry of Neural Network Quantization -- A Short Equivalence Proof of GPTQ and Babai's Algorithm
von: Birnick, Johann
Veröffentlicht: (2025)
von: Birnick, Johann
Veröffentlicht: (2025)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
von: Ha, SeungBum, et al.
Veröffentlicht: (2025)
von: Ha, SeungBum, et al.
Veröffentlicht: (2025)
Residual Reservoir Memory Networks
von: Pinna, Matteo, et al.
Veröffentlicht: (2025)
von: Pinna, Matteo, et al.
Veröffentlicht: (2025)
FreRA: A Frequency-Refined Augmentation for Contrastive Learning on Time Series Classification
von: Tian, Tian, et al.
Veröffentlicht: (2025)
von: Tian, Tian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
von: Rawal, Kaivalya, et al.
Veröffentlicht: (2026) -
Choosing DAG Models Using Markov and Minimal Edge Count in the Absence of Ground Truth
von: Ramsey, Joseph D., et al.
Veröffentlicht: (2024) -
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
von: Ribeiro, José, et al.
Veröffentlicht: (2022) -
How Reliable and Stable are Explanations of XAI Methods?
von: Ribeiro, José, et al.
Veröffentlicht: (2024) -
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
von: Gururaj, Shreyas, et al.
Veröffentlicht: (2025)