Beyond Output Faithfulness: Learning Attributions that Preserve Computational Pathways
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Siyu, Mcmillan, Kenneth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ABE: A Unified Framework for Robust and Faithful Attribution-Based Explainability
di: Zhu, Zhiyu, et al.
Pubblicazione: (2025)
di: Zhu, Zhiyu, et al.
Pubblicazione: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
di: Zhao, Zhixue, et al.
Pubblicazione: (2023)
di: Zhao, Zhixue, et al.
Pubblicazione: (2023)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
di: Huang, Xin, et al.
Pubblicazione: (2026)
di: Huang, Xin, et al.
Pubblicazione: (2026)
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
di: Su, Yiheng, et al.
Pubblicazione: (2023)
di: Su, Yiheng, et al.
Pubblicazione: (2023)
Faithful Density-Peaks Clustering via Matrix Computations on MPI Parallelization System
di: Xu, Ji, et al.
Pubblicazione: (2024)
di: Xu, Ji, et al.
Pubblicazione: (2024)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
di: Afolabi, Halimat, et al.
Pubblicazione: (2026)
di: Afolabi, Halimat, et al.
Pubblicazione: (2026)
Is Epistemic Uncertainty Faithfully Represented by Evidential Deep Learning Methods?
di: Jürgens, Mira, et al.
Pubblicazione: (2024)
di: Jürgens, Mira, et al.
Pubblicazione: (2024)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
di: Nur'aini, Khumaisa, et al.
Pubblicazione: (2026)
di: Nur'aini, Khumaisa, et al.
Pubblicazione: (2026)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
di: Guo, Yuhan, et al.
Pubblicazione: (2025)
di: Guo, Yuhan, et al.
Pubblicazione: (2025)
RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners
di: Gajjar, Jugal, et al.
Pubblicazione: (2026)
di: Gajjar, Jugal, et al.
Pubblicazione: (2026)
Learning Order Forest for Qualitative-Attribute Data Clustering
di: Zhao, Mingjie, et al.
Pubblicazione: (2026)
di: Zhao, Mingjie, et al.
Pubblicazione: (2026)
Reverse N-Wise Output-Oriented Testing for AI/ML and Quantum Computing Systems
di: Rihani, Lamine
Pubblicazione: (2026)
di: Rihani, Lamine
Pubblicazione: (2026)
Joint Input and Output Coordination for Class-Incremental Learning
di: Wang, Shuai, et al.
Pubblicazione: (2024)
di: Wang, Shuai, et al.
Pubblicazione: (2024)
Faithful Interpretation for Graph Neural Networks
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Learning Unified Distance Metric for Heterogeneous Attribute Data Clustering
di: Zhang, Yiqun, et al.
Pubblicazione: (2026)
di: Zhang, Yiqun, et al.
Pubblicazione: (2026)
RAGFormer: Learning Semantic Attributes and Topological Structure for Fraud Detection
di: Li, Haolin, et al.
Pubblicazione: (2024)
di: Li, Haolin, et al.
Pubblicazione: (2024)
Towards Metric-Faithful Neural Graph Matching
di: Shivottam, Jyotirmaya, et al.
Pubblicazione: (2026)
di: Shivottam, Jyotirmaya, et al.
Pubblicazione: (2026)
Beyond Gaussian Initializations: Signal Preserving Weight Initialization for Odd-Sigmoid Activations
di: Lee, Hyunwoo, et al.
Pubblicazione: (2025)
di: Lee, Hyunwoo, et al.
Pubblicazione: (2025)
Privacy-Preserving Federated Learning with Differentially Private Hyperdimensional Computing
di: Piran, Fardin Jalil, et al.
Pubblicazione: (2024)
di: Piran, Fardin Jalil, et al.
Pubblicazione: (2024)
On the Properties of Feature Attribution for Supervised Contrastive Learning
di: Arrighi, Leonardo, et al.
Pubblicazione: (2026)
di: Arrighi, Leonardo, et al.
Pubblicazione: (2026)
FaithLM: Towards Faithful Explanations for Large Language Models
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
Adjusting the Output of Decision Transformer with Action Gradient
di: Lin, Rui, et al.
Pubblicazione: (2025)
di: Lin, Rui, et al.
Pubblicazione: (2025)
Entropy-Preserving Reinforcement Learning
di: Petrenko, Aleksei, et al.
Pubblicazione: (2026)
di: Petrenko, Aleksei, et al.
Pubblicazione: (2026)
An Analysis under a Unified Fomulation of Learning Algorithms with Output Constraints
di: Song, Mooho, et al.
Pubblicazione: (2024)
di: Song, Mooho, et al.
Pubblicazione: (2024)
AL-GNN: Privacy-Preserving and Replay-Free Continual Graph Learning via Analytic Learning
di: Zhang, Xuling, et al.
Pubblicazione: (2025)
di: Zhang, Xuling, et al.
Pubblicazione: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
di: Yan, Ge, et al.
Pubblicazione: (2025)
di: Yan, Ge, et al.
Pubblicazione: (2025)
Faithful and Fast Influence Function via Advanced Sampling
di: Koh, Jungyeon, et al.
Pubblicazione: (2025)
di: Koh, Jungyeon, et al.
Pubblicazione: (2025)
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
di: DatologyAI, et al.
Pubblicazione: (2026)
di: DatologyAI, et al.
Pubblicazione: (2026)
LIMEtree: Consistent and Faithful Surrogate Explanations of Multiple Classes
di: Sokol, Kacper, et al.
Pubblicazione: (2020)
di: Sokol, Kacper, et al.
Pubblicazione: (2020)
Faithful Differentiable Reasoning with Reshuffled Region-based Embeddings
di: Pavlovic, Aleksandar, et al.
Pubblicazione: (2024)
di: Pavlovic, Aleksandar, et al.
Pubblicazione: (2024)
Towards Faithful Explanations: Boosting Rationalization with Shortcuts Discovery
di: Yue, Linan, et al.
Pubblicazione: (2024)
di: Yue, Linan, et al.
Pubblicazione: (2024)
Derivation of Output Correlation Inferences for Multi-Output (aka Multi-Task) Gaussian Process
di: Watanabe, Shuhei
Pubblicazione: (2025)
di: Watanabe, Shuhei
Pubblicazione: (2025)
Preserve Support, Not Correspondence: Dynamic Routing for Offline Reinforcement Learning
di: Mu, Zhancun, et al.
Pubblicazione: (2026)
di: Mu, Zhancun, et al.
Pubblicazione: (2026)
Semantic-Inductive Attribute Selection for Zero-Shot Learning
di: Herrera-Aranda, Juan Jose, et al.
Pubblicazione: (2025)
di: Herrera-Aranda, Juan Jose, et al.
Pubblicazione: (2025)
VRAIL: Vectorized Reward-based Attribution for Interpretable Learning
di: Kim, Jina, et al.
Pubblicazione: (2025)
di: Kim, Jina, et al.
Pubblicazione: (2025)
ANML: Attribution-Native Machine Learning with Guaranteed Robustness
di: Zahn, Oliver, et al.
Pubblicazione: (2026)
di: Zahn, Oliver, et al.
Pubblicazione: (2026)
Generalized Group Data Attribution
di: Ley, Dan, et al.
Pubblicazione: (2024)
di: Ley, Dan, et al.
Pubblicazione: (2024)
Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning
di: Lillo, Lute, et al.
Pubblicazione: (2026)
di: Lillo, Lute, et al.
Pubblicazione: (2026)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
di: Swaroop, Anand, et al.
Pubblicazione: (2025)
di: Swaroop, Anand, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ABE: A Unified Framework for Robust and Faithful Attribution-Based Explainability
di: Zhu, Zhiyu, et al.
Pubblicazione: (2025) -
Incorporating Attribution Importance for Improving Faithfulness Metrics
di: Zhao, Zhixue, et al.
Pubblicazione: (2023) -
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
di: Huang, Xin, et al.
Pubblicazione: (2026) -
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
di: Su, Yiheng, et al.
Pubblicazione: (2023) -
Faithful Density-Peaks Clustering via Matrix Computations on MPI Parallelization System
di: Xu, Ji, et al.
Pubblicazione: (2024)