Fusing Rewards and Preferences in Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Khorasani, Sadegh, Salehkaleybar, Saber, Kiyavash, Negar, Grossglauser, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hierarchical Reinforcement Learning with Targeted Causal Interventions
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Inference Time Causal Probing in LLMs
by: Khorasani, Sadegh, et al.
Published: (2026)
by: Khorasani, Sadegh, et al.
Published: (2026)
Efficiently Escaping Saddle Points for Policy Optimization
by: Khorasani, Sadegh, et al.
Published: (2023)
by: Khorasani, Sadegh, et al.
Published: (2023)
Measuring IIA Violations in Similarity Choices with Bayesian Models
by: Corrêa, Hugo Sales, et al.
Published: (2025)
by: Corrêa, Hugo Sales, et al.
Published: (2025)
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
by: Quadros, André, et al.
Published: (2025)
by: Quadros, André, et al.
Published: (2025)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Perfecting Aircraft Maneuvers with Reinforcement Learning
by: Cilan, Atahan, et al.
Published: (2026)
by: Cilan, Atahan, et al.
Published: (2026)
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
by: Li, Bangzheng, et al.
Published: (2024)
by: Li, Bangzheng, et al.
Published: (2024)
Cost and Reward Infused Metric Elicitation
by: Bhateja, Chethan, et al.
Published: (2025)
by: Bhateja, Chethan, et al.
Published: (2025)
Action-Dependent Optimality-Preserving Reward Shaping
by: Forbes, Grant C., et al.
Published: (2025)
by: Forbes, Grant C., et al.
Published: (2025)
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Data-Incremental Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2024)
by: Gai, Sibo, et al.
Published: (2024)
A Survey of Reinforcement Learning from Human Feedback
by: Kaufmann, Timo, et al.
Published: (2023)
by: Kaufmann, Timo, et al.
Published: (2023)
OER: Offline Experience Replay for Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2023)
by: Gai, Sibo, et al.
Published: (2023)
A Comparative Analysis of Reinforcement Learning and Conventional Deep Learning Approaches for Bearing Fault Diagnosis
by: Çakır, Efe, et al.
Published: (2025)
by: Çakır, Efe, et al.
Published: (2025)
Safe Reinforcement Learning with Preference-based Constraint Inference
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Fractional Policy Gradients: Reinforcement Learning with Long-Term Memory
by: Pawar, Urvi, et al.
Published: (2025)
by: Pawar, Urvi, et al.
Published: (2025)
Playing Hex and Counter Wargames using Reinforcement Learning and Recurrent Neural Networks
by: Palma, Guilherme, et al.
Published: (2025)
by: Palma, Guilherme, et al.
Published: (2025)
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Distributional Reinforcement Learning for Condition-Based Maintenance of Multi-Pump Equipment
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
ES-C51: Expected Sarsa Based C51 Distributional Reinforcement Learning Algorithm
by: Tandon, Rijul, et al.
Published: (2025)
by: Tandon, Rijul, et al.
Published: (2025)
Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
by: Chapman, James, et al.
Published: (2025)
by: Chapman, James, et al.
Published: (2025)
From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
by: Tomashevskiy, Timofey
Published: (2026)
by: Tomashevskiy, Timofey
Published: (2026)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
by: Chen, Wen-Tse, et al.
Published: (2024)
by: Chen, Wen-Tse, et al.
Published: (2024)
FlowRL: Flow-Augmented Few-Shot Reinforcement Learning for Semi-Structured Sensor Data
by: Pivezhandi, Mohammad, et al.
Published: (2024)
by: Pivezhandi, Mohammad, et al.
Published: (2024)
SQARL: A Size-Agnostic Reinforcement Learning approach for Circuit Allocation in Distributed Quantum Architectures
by: Carballo, Víctor, et al.
Published: (2026)
by: Carballo, Víctor, et al.
Published: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026)
by: Tang, Wenjie, et al.
Published: (2026)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
by: Hisaki, Yukinari, et al.
Published: (2024)
by: Hisaki, Yukinari, et al.
Published: (2024)
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
by: Mongaras, Gabriel, et al.
Published: (2026)
by: Mongaras, Gabriel, et al.
Published: (2026)
I-GLIDE: Input Groups for Latent Health Indicators in Degradation Estimation
by: Thil, Lucas, et al.
Published: (2025)
by: Thil, Lucas, et al.
Published: (2025)
Learning Transferable Predictability Representations
by: Goswami, Diyali, et al.
Published: (2026)
by: Goswami, Diyali, et al.
Published: (2026)
Symmetric Equilibrium Learning of VAEs
by: Flach, Boris, et al.
Published: (2023)
by: Flach, Boris, et al.
Published: (2023)
FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models
by: Polly, Fabien
Published: (2026)
by: Polly, Fabien
Published: (2026)
Intervening to Learn and Compose Causally Disentangled Representations
by: Markham, Alex, et al.
Published: (2025)
by: Markham, Alex, et al.
Published: (2025)
Boosting gets full Attention for Relational Learning
by: Guillame-Bert, Mathieu, et al.
Published: (2024)
by: Guillame-Bert, Mathieu, et al.
Published: (2024)
ArrowFlow: Hierarchical Machine Learning in the Space of Permutations
by: Yilmaz, Ozgur
Published: (2026)
by: Yilmaz, Ozgur
Published: (2026)
Trusted Multi-view Learning under Noisy Supervision
by: Zhang, Yilin, et al.
Published: (2024)
by: Zhang, Yilin, et al.
Published: (2024)
Empowering Graph Invariance Learning with Deep Spurious Infomax
by: Yao, Tianjun, et al.
Published: (2024)
by: Yao, Tianjun, et al.
Published: (2024)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
by: Sun, Yuhui, et al.
Published: (2025)
by: Sun, Yuhui, et al.
Published: (2025)
Similar Items
-
Hierarchical Reinforcement Learning with Targeted Causal Interventions
by: Khorasani, Sadegh, et al.
Published: (2025) -
Inference Time Causal Probing in LLMs
by: Khorasani, Sadegh, et al.
Published: (2026) -
Efficiently Escaping Saddle Points for Policy Optimization
by: Khorasani, Sadegh, et al.
Published: (2023) -
Measuring IIA Violations in Similarity Choices with Bayesian Models
by: Corrêa, Hugo Sales, et al.
Published: (2025) -
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
by: Quadros, André, et al.
Published: (2025)