Target-Aligned Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Pleiss, Leonard S., Harrison, James, Schiffer, Maximilian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic Monitoring Environments for Reinforcement Learning
by: Pleiss, Leonard, et al.
Published: (2026)
by: Pleiss, Leonard, et al.
Published: (2026)
Reliability-Adjusted Prioritized Experience Replay
by: Pleiss, Leonard S., et al.
Published: (2025)
by: Pleiss, Leonard S., et al.
Published: (2025)
Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces
by: Hoppe, Heiko, et al.
Published: (2026)
by: Hoppe, Heiko, et al.
Published: (2026)
Disentangling generalization and memorization in large language models using chess
by: Pleiss, Leonard S., et al.
Published: (2026)
by: Pleiss, Leonard S., et al.
Published: (2026)
Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning
by: Kim, Minung, et al.
Published: (2026)
by: Kim, Minung, et al.
Published: (2026)
Neural Cluster First, Route Second: One-Shot Capacitated Vehicle Routing via Differentiable Optimal Transport
by: Chin, Samuel J. K., et al.
Published: (2026)
by: Chin, Samuel J. K., et al.
Published: (2026)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
by: Cinquin, Tristan, et al.
Published: (2025)
by: Cinquin, Tristan, et al.
Published: (2025)
Risk-Sensitive Soft Actor-Critic for Robust Deep Reinforcement Learning under Distribution Shifts
by: Enders, Tobias, et al.
Published: (2024)
by: Enders, Tobias, et al.
Published: (2024)
Preference-aware compensation policies for crowdsourced on-demand services
by: Nouli, Georgina, et al.
Published: (2025)
by: Nouli, Georgina, et al.
Published: (2025)
Dynamic Neighborhood Construction for Structured Large Discrete Action Spaces
by: Akkerman, Fabian, et al.
Published: (2023)
by: Akkerman, Fabian, et al.
Published: (2023)
LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment
by: Li, Shipeng, et al.
Published: (2025)
by: Li, Shipeng, et al.
Published: (2025)
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
by: Yang, Ningyuan, et al.
Published: (2026)
by: Yang, Ningyuan, et al.
Published: (2026)
Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2025)
by: Vincent, Théo, et al.
Published: (2025)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
by: Nguyen, Thanh Thi, et al.
Published: (2025)
by: Nguyen, Thanh Thi, et al.
Published: (2025)
Evaluating Reinforcement Learning Algorithms for Navigation in Simulated Robotic Quadrupeds: A Comparative Study Inspired by Guide Dog Behaviour
by: Harrison, Emma M. A.
Published: (2025)
by: Harrison, Emma M. A.
Published: (2025)
A Survey of Explainable Reinforcement Learning: Targets, Methods and Needs
by: Saulières, Léo
Published: (2025)
by: Saulières, Léo
Published: (2025)
Aligning Findings with Diagnosis: A Self-Consistent Reinforcement Learning Framework for Trustworthy Radiology Reporting
by: Zhao, Kun, et al.
Published: (2026)
by: Zhao, Kun, et al.
Published: (2026)
Adaptive $Q$-Network: On-the-fly Target Selection for Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
Surrogate Fitness Metrics for Interpretable Reinforcement Learning
by: Altmann, Philipp, et al.
Published: (2025)
by: Altmann, Philipp, et al.
Published: (2025)
Causally Aligned Curriculum Learning
by: Li, Mingxuan, et al.
Published: (2025)
by: Li, Mingxuan, et al.
Published: (2025)
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
by: Zhang, Xinnan, et al.
Published: (2025)
by: Zhang, Xinnan, et al.
Published: (2025)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
by: Ma, Wenhan, et al.
Published: (2025)
by: Ma, Wenhan, et al.
Published: (2025)
Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation
by: Bai, Fengshuo, et al.
Published: (2024)
by: Bai, Fengshuo, et al.
Published: (2024)
Constrained Latent Action Policies for Model-Based Offline Reinforcement Learning
by: Alles, Marvin, et al.
Published: (2024)
by: Alles, Marvin, et al.
Published: (2024)
Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits
by: Mahlau, Yannik, et al.
Published: (2025)
by: Mahlau, Yannik, et al.
Published: (2025)
Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions
by: Abate, Alessandro, et al.
Published: (2026)
by: Abate, Alessandro, et al.
Published: (2026)
On Distributional Reinforcement Learning in Chaotic Dynamical Systems
by: Rudd-Jones, James, et al.
Published: (2026)
by: Rudd-Jones, James, et al.
Published: (2026)
Universal Neural Functionals
by: Zhou, Allan, et al.
Published: (2024)
by: Zhou, Allan, et al.
Published: (2024)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025)
by: Dabas, Mahavir, et al.
Published: (2025)
Research and Design on Intelligent Recognition of Unordered Targets for Robots Based on Reinforcement Learning
by: Mao, Yiting, et al.
Published: (2025)
by: Mao, Yiting, et al.
Published: (2025)
Structure-Aligned Protein Language Model
by: Chen, Can, et al.
Published: (2025)
by: Chen, Can, et al.
Published: (2025)
HUGO -- Highlighting Unseen Grid Options: Combining Deep Reinforcement Learning with a Heuristic Target Topology Approach
by: Lehna, Malte, et al.
Published: (2024)
by: Lehna, Malte, et al.
Published: (2024)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
Early Prediction of Sepsis: Feature-Aligned Transfer Learning
by: Komolafe, Oyindolapo O., et al.
Published: (2025)
by: Komolafe, Oyindolapo O., et al.
Published: (2025)
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
by: Geles, Ismail, et al.
Published: (2026)
by: Geles, Ismail, et al.
Published: (2026)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024)
by: Tan, Weihao, et al.
Published: (2024)
Privileged Sensing Scaffolds Reinforcement Learning
by: Hu, Edward S., et al.
Published: (2024)
by: Hu, Edward S., et al.
Published: (2024)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Similar Items
-
Synthetic Monitoring Environments for Reinforcement Learning
by: Pleiss, Leonard, et al.
Published: (2026) -
Reliability-Adjusted Prioritized Experience Replay
by: Pleiss, Leonard S., et al.
Published: (2025) -
Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces
by: Hoppe, Heiko, et al.
Published: (2026) -
Disentangling generalization and memorization in large language models using chess
by: Pleiss, Leonard S., et al.
Published: (2026) -
Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning
by: Kim, Minung, et al.
Published: (2026)