Can Differentiable Decision Trees Enable Interpretable Reward Learning from Human Feedback?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kalra, Akansha, Brown, Daniel S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoCP-GNN: Robust Conformal Prediction for Graph Neural Networks in Node-Classification
von: Akansha, S.
Veröffentlicht: (2024)
von: Akansha, S.
Veröffentlicht: (2024)
Adaptive Querying for Reward Learning from Human Feedback
von: Anand, Yashwanthi, et al.
Veröffentlicht: (2024)
von: Anand, Yashwanthi, et al.
Veröffentlicht: (2024)
Conditional Shift-Robust Conformal Prediction for Graph Neural Network
von: Akansha, S.
Veröffentlicht: (2024)
von: Akansha, S.
Veröffentlicht: (2024)
Reward Learning from Multiple Feedback Types
von: Metz, Yannick, et al.
Veröffentlicht: (2025)
von: Metz, Yannick, et al.
Veröffentlicht: (2025)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
von: Kim, Kihyun, et al.
Veröffentlicht: (2024)
von: Kim, Kihyun, et al.
Veröffentlicht: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
von: Blaser, Ethan, et al.
Veröffentlicht: (2026)
von: Blaser, Ethan, et al.
Veröffentlicht: (2026)
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
von: Hwang, Minjune, et al.
Veröffentlicht: (2026)
von: Hwang, Minjune, et al.
Veröffentlicht: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
von: Zhang, Shun, et al.
Veröffentlicht: (2024)
von: Zhang, Shun, et al.
Veröffentlicht: (2024)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2025)
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2025)
An Interpretable Client Decision Tree Aggregation process for Federated Learning
von: Argente-Garrido, Alberto, et al.
Veröffentlicht: (2024)
von: Argente-Garrido, Alberto, et al.
Veröffentlicht: (2024)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
VRAIL: Vectorized Reward-based Attribution for Interpretable Learning
von: Kim, Jina, et al.
Veröffentlicht: (2025)
von: Kim, Jina, et al.
Veröffentlicht: (2025)
What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning?
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
von: Karimi, Zohre, et al.
Veröffentlicht: (2024)
von: Karimi, Zohre, et al.
Veröffentlicht: (2024)
Multi-Task Reward Learning from Human Ratings
von: Wu, Mingkang, et al.
Veröffentlicht: (2025)
von: Wu, Mingkang, et al.
Veröffentlicht: (2025)
Decision-Focused Model-based Reinforcement Learning for Reward Transfer
von: Sharma, Abhishek, et al.
Veröffentlicht: (2023)
von: Sharma, Abhishek, et al.
Veröffentlicht: (2023)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
von: Baur, Raphaël, et al.
Veröffentlicht: (2026)
von: Baur, Raphaël, et al.
Veröffentlicht: (2026)
Approximation-Free Differentiable Oblique Decision Trees
von: Panda, Subrat Prasad, et al.
Veröffentlicht: (2026)
von: Panda, Subrat Prasad, et al.
Veröffentlicht: (2026)
Fusing Reward and Dueling Feedback in Stochastic Bandits
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
von: Zheng, Qinqing, et al.
Veröffentlicht: (2024)
von: Zheng, Qinqing, et al.
Veröffentlicht: (2024)
Decision Predicate Graphs: Enhancing Interpretability in Tree Ensembles
von: Arrighi, Leonardo, et al.
Veröffentlicht: (2024)
von: Arrighi, Leonardo, et al.
Veröffentlicht: (2024)
Contrastive Preference Learning: Learning from Human Feedback without RL
von: Hejna, Joey, et al.
Veröffentlicht: (2023)
von: Hejna, Joey, et al.
Veröffentlicht: (2023)
How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies
von: Kalra, Akansha, et al.
Veröffentlicht: (2025)
von: Kalra, Akansha, et al.
Veröffentlicht: (2025)
Understanding the Learning Dynamics of Alignment with Human Feedback
von: Im, Shawn, et al.
Veröffentlicht: (2024)
von: Im, Shawn, et al.
Veröffentlicht: (2024)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
von: Vora, Kevin, et al.
Veröffentlicht: (2025)
von: Vora, Kevin, et al.
Veröffentlicht: (2025)
Reward Design for Justifiable Sequential Decision-Making
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
von: Kim, Gihoon, et al.
Veröffentlicht: (2026)
von: Kim, Gihoon, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Symbolic Reward Machines
von: Krug, Thomas, et al.
Veröffentlicht: (2026)
von: Krug, Thomas, et al.
Veröffentlicht: (2026)
Towards Adaptive Deep Learning: Model Elasticity via Prune-and-Grow CNN Architectures
von: Mangal, Pooja, et al.
Veröffentlicht: (2025)
von: Mangal, Pooja, et al.
Veröffentlicht: (2025)
Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
von: Trapasso, Alessandro, et al.
Veröffentlicht: (2025)
von: Trapasso, Alessandro, et al.
Veröffentlicht: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
Learning Human-like Representations to Enable Learning Human Values
von: Wynn, Andrea, et al.
Veröffentlicht: (2023)
von: Wynn, Andrea, et al.
Veröffentlicht: (2023)
Adaptive Preference Scaling for Reinforcement Learning with Human Feedback
von: Hong, Ilgee, et al.
Veröffentlicht: (2024)
von: Hong, Ilgee, et al.
Veröffentlicht: (2024)
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Batch Active Learning of Reward Functions from Human Preferences
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RoCP-GNN: Robust Conformal Prediction for Graph Neural Networks in Node-Classification
von: Akansha, S.
Veröffentlicht: (2024) -
Adaptive Querying for Reward Learning from Human Feedback
von: Anand, Yashwanthi, et al.
Veröffentlicht: (2024) -
Conditional Shift-Robust Conformal Prediction for Graph Neural Network
von: Akansha, S.
Veröffentlicht: (2024) -
Reward Learning from Multiple Feedback Types
von: Metz, Yannick, et al.
Veröffentlicht: (2025) -
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
von: Luo, Renjie, et al.
Veröffentlicht: (2025)