Reducing Reward Dependence in RL Through Adaptive Confidence Discounting
Fuente:
arXiv
Guardado en:
| Autores principales: | Satici, Muhammed Yusuf, Roberts, David L. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Autonomous Curriculum Design via Relative Entropy Based Task Modifications
por: Satici, Muhammed Yusuf, et al.
Publicado: (2025)
por: Satici, Muhammed Yusuf, et al.
Publicado: (2025)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
por: Grand-Clément, Julien, et al.
Publicado: (2023)
por: Grand-Clément, Julien, et al.
Publicado: (2023)
Discounted Adaptive Online Learning: Towards Better Regularization
por: Zhang, Zhiyu, et al.
Publicado: (2024)
por: Zhang, Zhiyu, et al.
Publicado: (2024)
Action-Dependent Optimality-Preserving Reward Shaping
por: Forbes, Grant C., et al.
Publicado: (2025)
por: Forbes, Grant C., et al.
Publicado: (2025)
Analyzing and Bridging the Gap between Maximizing Total Reward and Discounted Reward in Deep Reinforcement Learning
por: Yin, Shuyu, et al.
Publicado: (2024)
por: Yin, Shuyu, et al.
Publicado: (2024)
Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
por: Hong, Kihyuk, et al.
Publicado: (2024)
por: Hong, Kihyuk, et al.
Publicado: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
por: Singha, Disha
Publicado: (2026)
por: Singha, Disha
Publicado: (2026)
The Impact of Post-training on Data Contamination
por: Kocyigit, Muhammed Yusuf, et al.
Publicado: (2026)
por: Kocyigit, Muhammed Yusuf, et al.
Publicado: (2026)
Adaptive Discounting of Training Time Attacks
por: Bector, Ridhima, et al.
Publicado: (2024)
por: Bector, Ridhima, et al.
Publicado: (2024)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
por: Muslimani, Calarina, et al.
Publicado: (2025)
por: Muslimani, Calarina, et al.
Publicado: (2025)
AdaGamma: State-Dependent Discounting for Temporal Adaptation in Reinforcement Learning
por: Wang, Yaomin, et al.
Publicado: (2026)
por: Wang, Yaomin, et al.
Publicado: (2026)
Reward Compatibility: A Framework for Inverse RL
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
por: Zurek, Matthew, et al.
Publicado: (2024)
por: Zurek, Matthew, et al.
Publicado: (2024)
On-Policy RL with Optimal Reward Baseline
por: Hao, Yaru, et al.
Publicado: (2025)
por: Hao, Yaru, et al.
Publicado: (2025)
Online Recommendations for Agents with Discounted Adaptive Preferences
por: Agarwal, Arpit, et al.
Publicado: (2023)
por: Agarwal, Arpit, et al.
Publicado: (2023)
Natural Policy Gradient for Average Reward Non-Stationary RL
por: Jali, Neharika, et al.
Publicado: (2025)
por: Jali, Neharika, et al.
Publicado: (2025)
Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting
por: Feng, Yunzhen, et al.
Publicado: (2025)
por: Feng, Yunzhen, et al.
Publicado: (2025)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
por: Cai, Xin-Qiang, et al.
Publicado: (2026)
por: Cai, Xin-Qiang, et al.
Publicado: (2026)
Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs
por: Thoppe, Gugan, et al.
Publicado: (2026)
por: Thoppe, Gugan, et al.
Publicado: (2026)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
RL + Transformer = A General-Purpose Problem Solver
por: Rentschler, Micah, et al.
Publicado: (2025)
por: Rentschler, Micah, et al.
Publicado: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
por: Tan, Chuyi, et al.
Publicado: (2025)
por: Tan, Chuyi, et al.
Publicado: (2025)
Gatekeeper: Improving Model Cascades Through Confidence Tuning
por: Rabanser, Stephan, et al.
Publicado: (2025)
por: Rabanser, Stephan, et al.
Publicado: (2025)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
por: Cheng, Tianhao, et al.
Publicado: (2026)
por: Cheng, Tianhao, et al.
Publicado: (2026)
Less is more? Rewards in RL for Cyber Defence
por: Bates, Elizabeth, et al.
Publicado: (2025)
por: Bates, Elizabeth, et al.
Publicado: (2025)
Learning to Reason Efficiently with Discounted Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
por: Jin, Can, et al.
Publicado: (2025)
por: Jin, Can, et al.
Publicado: (2025)
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
por: Mukherjee, Subhojyoti, et al.
Publicado: (2025)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2025)
General Flexible $f$-divergence for Challenging Offline RL Datasets with Low Stochasticity and Diverse Behavior Policies
por: Wang, Jianxun, et al.
Publicado: (2026)
por: Wang, Jianxun, et al.
Publicado: (2026)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
por: Zhang, Bonan, et al.
Publicado: (2025)
por: Zhang, Bonan, et al.
Publicado: (2025)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
por: Kim, Kyuyoung, et al.
Publicado: (2024)
por: Kim, Kyuyoung, et al.
Publicado: (2024)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
por: Salimi, Moein, et al.
Publicado: (2026)
por: Salimi, Moein, et al.
Publicado: (2026)
Reducing Oracle Feedback with Vision-Language Embeddings for Preference-Based RL
por: Ghosh, Udita, et al.
Publicado: (2026)
por: Ghosh, Udita, et al.
Publicado: (2026)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
por: Mahankali, Arvind, et al.
Publicado: (2026)
por: Mahankali, Arvind, et al.
Publicado: (2026)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
por: Lyu, Lixing, et al.
Publicado: (2025)
por: Lyu, Lixing, et al.
Publicado: (2025)
Bridging the Gap Between Average and Discounted TD Learning
por: Tian, Haoxing, et al.
Publicado: (2026)
por: Tian, Haoxing, et al.
Publicado: (2026)
Online Linear Regression in Dynamic Environments via Discounting
por: Jacobsen, Andrew, et al.
Publicado: (2024)
por: Jacobsen, Andrew, et al.
Publicado: (2024)
When Your LLM Reaches End-of-Life: A Framework for Confident Model Migration in Production Systems
por: Casey, Emma, et al.
Publicado: (2026)
por: Casey, Emma, et al.
Publicado: (2026)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
por: Mondal, Washim Uddin, et al.
Publicado: (2023)
por: Mondal, Washim Uddin, et al.
Publicado: (2023)
Ejemplares similares
-
Autonomous Curriculum Design via Relative Entropy Based Task Modifications
por: Satici, Muhammed Yusuf, et al.
Publicado: (2025) -
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
por: Grand-Clément, Julien, et al.
Publicado: (2023) -
Discounted Adaptive Online Learning: Towards Better Regularization
por: Zhang, Zhiyu, et al.
Publicado: (2024) -
Action-Dependent Optimality-Preserving Reward Shaping
por: Forbes, Grant C., et al.
Publicado: (2025) -
Analyzing and Bridging the Gap between Maximizing Total Reward and Discounted Reward in Deep Reinforcement Learning
por: Yin, Shuyu, et al.
Publicado: (2024)