VRAIL: Vectorized Reward-based Attribution for Interpretable Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Jina, Jang, Youjin, Han, Jeongjin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Teaching AI to Remember: Insights from Brain-Inspired Replay in Continual Learning
von: Kim, Jina
Veröffentlicht: (2025)
von: Kim, Jina
Veröffentlicht: (2025)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
von: Kim, Yoonjeon, et al.
Veröffentlicht: (2025)
von: Kim, Yoonjeon, et al.
Veröffentlicht: (2025)
What Is the Point of Equality in Machine Learning Fairness? Beyond Equality of Opportunity
von: Kong, Youjin
Veröffentlicht: (2025)
von: Kong, Youjin
Veröffentlicht: (2025)
Learning to Price: Interpretable Attribute-Level Models for Dynamic Markets
von: Sethuraman, Srividhya, et al.
Veröffentlicht: (2026)
von: Sethuraman, Srividhya, et al.
Veröffentlicht: (2026)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
AIM: Attributing, Interpreting, Mitigating Data Unfairness
von: Liu, Zhining, et al.
Veröffentlicht: (2024)
von: Liu, Zhining, et al.
Veröffentlicht: (2024)
Attributions All the Way Down? The Metagame of Interpretability
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
Interpretable Deep Learning for Stock Returns: A Consensus-Bottleneck Asset Pricing Model
von: Kim, Changeun, et al.
Veröffentlicht: (2025)
von: Kim, Changeun, et al.
Veröffentlicht: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
Can Differentiable Decision Trees Enable Interpretable Reward Learning from Human Feedback?
von: Kalra, Akansha, et al.
Veröffentlicht: (2023)
von: Kalra, Akansha, et al.
Veröffentlicht: (2023)
Interpreting Language Reward Models via Contrastive Explanations
von: Jiang, Junqi, et al.
Veröffentlicht: (2024)
von: Jiang, Junqi, et al.
Veröffentlicht: (2024)
Robust Molecular Property Prediction via Densifying Scarce Labeled Data
von: Kim, Jina, et al.
Veröffentlicht: (2025)
von: Kim, Jina, et al.
Veröffentlicht: (2025)
Enhancing Model Interpretability with Local Attribution over Global Exploration
von: Zhu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Zhu, Zhiyu, et al.
Veröffentlicht: (2024)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
reward-lens: A Mechanistic Interpretability Library for Reward Models
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
Rectifying Shortcut Behaviors in Preference-based Reward Learning
von: Ye, Wenqian, et al.
Veröffentlicht: (2025)
von: Ye, Wenqian, et al.
Veröffentlicht: (2025)
Auxiliary Reward Generation with Transition Distance Representation Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
Interpretable Prototype-based Graph Information Bottleneck
von: Seo, Sangwoo, et al.
Veröffentlicht: (2023)
von: Seo, Sangwoo, et al.
Veröffentlicht: (2023)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
von: Jang, Sooyoung, et al.
Veröffentlicht: (2021)
von: Jang, Sooyoung, et al.
Veröffentlicht: (2021)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
von: Liu, Dengcan, et al.
Veröffentlicht: (2026)
von: Liu, Dengcan, et al.
Veröffentlicht: (2026)
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
von: Pepper, Keenan, et al.
Veröffentlicht: (2026)
von: Pepper, Keenan, et al.
Veröffentlicht: (2026)
Decision-Focused Model-based Reinforcement Learning for Reward Transfer
von: Sharma, Abhishek, et al.
Veröffentlicht: (2023)
von: Sharma, Abhishek, et al.
Veröffentlicht: (2023)
Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
von: Okudo, Takato, et al.
Veröffentlicht: (2021)
von: Okudo, Takato, et al.
Veröffentlicht: (2021)
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
von: Choi, Heewoong, et al.
Veröffentlicht: (2024)
von: Choi, Heewoong, et al.
Veröffentlicht: (2024)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
von: Rajaram, Sara, et al.
Veröffentlicht: (2025)
von: Rajaram, Sara, et al.
Veröffentlicht: (2025)
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
von: Metcalf, Katherine, et al.
Veröffentlicht: (2024)
von: Metcalf, Katherine, et al.
Veröffentlicht: (2024)
Residual Reward Models for Preference-based Reinforcement Learning
von: Cao, Chenyang, et al.
Veröffentlicht: (2025)
von: Cao, Chenyang, et al.
Veröffentlicht: (2025)
Unveiling the Significance of Toddler-Inspired Reward Transition in Goal-Oriented Reinforcement Learning
von: Park, Junseok, et al.
Veröffentlicht: (2024)
von: Park, Junseok, et al.
Veröffentlicht: (2024)
Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning
von: Yin, Junxi, et al.
Veröffentlicht: (2025)
von: Yin, Junxi, et al.
Veröffentlicht: (2025)
SemiReward: A General Reward Model for Semi-supervised Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2022)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2022)
Influence-based Attributions can be Manipulated
von: Yadav, Chhavi, et al.
Veröffentlicht: (2024)
von: Yadav, Chhavi, et al.
Veröffentlicht: (2024)
Impossibility Theorems for Feature Attribution
von: Bilodeau, Blair, et al.
Veröffentlicht: (2022)
von: Bilodeau, Blair, et al.
Veröffentlicht: (2022)
Towards Case-based Interpretability for Medical Federated Learning
von: Latorre, Laura, et al.
Veröffentlicht: (2024)
von: Latorre, Laura, et al.
Veröffentlicht: (2024)
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
von: Eaton, Kenneth, et al.
Veröffentlicht: (2024)
von: Eaton, Kenneth, et al.
Veröffentlicht: (2024)
OMG-RL:Offline Model-based Guided Reward Learning for Heparin Treatment
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Teaching AI to Remember: Insights from Brain-Inspired Replay in Continual Learning
von: Kim, Jina
Veröffentlicht: (2025) -
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024) -
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
von: Kim, Yoonjeon, et al.
Veröffentlicht: (2025) -
What Is the Point of Equality in Machine Learning Fairness? Beyond Equality of Opportunity
von: Kong, Youjin
Veröffentlicht: (2025) -
Learning to Price: Interpretable Attribute-Level Models for Dynamic Markets
von: Sethuraman, Srividhya, et al.
Veröffentlicht: (2026)