Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Yuxiao, Zhang, Weitong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How to Provably Improve Return Conditioned Supervised Learning?
por: Liu, Zhishuai, et al.
Publicado: (2025)
por: Liu, Zhishuai, et al.
Publicado: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
por: Yang, Yuxiao, et al.
Publicado: (2026)
por: Yang, Yuxiao, et al.
Publicado: (2026)
Imitation Learning as Return Distribution Matching
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
On the Diminishing Returns of Width for Continual Learning
por: Guha, Etash, et al.
Publicado: (2024)
por: Guha, Etash, et al.
Publicado: (2024)
STAS: Spatial-Temporal Return Decomposition for Multi-agent Reinforcement Learning
por: Chen, Sirui, et al.
Publicado: (2023)
por: Chen, Sirui, et al.
Publicado: (2023)
Meta-Learning Reinforcement Learning for Crypto-Return Prediction
por: Wang, Junqiao, et al.
Publicado: (2025)
por: Wang, Junqiao, et al.
Publicado: (2025)
In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning
por: Tu, Songjun, et al.
Publicado: (2024)
por: Tu, Songjun, et al.
Publicado: (2024)
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
por: Wang, Ruhan, et al.
Publicado: (2024)
por: Wang, Ruhan, et al.
Publicado: (2024)
Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
por: Lee, Dongkwan, et al.
Publicado: (2025)
por: Lee, Dongkwan, et al.
Publicado: (2025)
SACn: Soft Actor-Critic with n-step Returns
por: Łyskawa, Jakub, et al.
Publicado: (2025)
por: Łyskawa, Jakub, et al.
Publicado: (2025)
Beyond Expected Return: Accounting for Policy Reproducibility when Evaluating Reinforcement Learning Algorithms
por: Flageat, Manon, et al.
Publicado: (2023)
por: Flageat, Manon, et al.
Publicado: (2023)
Moments Matter:Stabilizing Policy Optimization using Return Distributions
por: Jabs, Dennis, et al.
Publicado: (2026)
por: Jabs, Dennis, et al.
Publicado: (2026)
Safe Online Bid Optimization with Return on Investment and Budget Constraints
por: Castiglioni, Matteo, et al.
Publicado: (2022)
por: Castiglioni, Matteo, et al.
Publicado: (2022)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
por: Goodall, Alexander W., et al.
Publicado: (2025)
por: Goodall, Alexander W., et al.
Publicado: (2025)
Multi-objective Reinforcement Learning with Nonlinear Preferences: Provable Approximation for Maximizing Expected Scalarized Return
por: Peng, Nianli, et al.
Publicado: (2023)
por: Peng, Nianli, et al.
Publicado: (2023)
Target Return Optimizer for Multi-Game Decision Transformer
por: Tatematsu, Kensuke, et al.
Publicado: (2025)
por: Tatematsu, Kensuke, et al.
Publicado: (2025)
More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning
por: Ma, Zhipeng, et al.
Publicado: (2024)
por: Ma, Zhipeng, et al.
Publicado: (2024)
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
por: Mead, Harry, et al.
Publicado: (2025)
por: Mead, Harry, et al.
Publicado: (2025)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
por: Kiyohara, Haruka, et al.
Publicado: (2023)
por: Kiyohara, Haruka, et al.
Publicado: (2023)
Optimizing Return Distributions with Distributional Dynamic Programming
por: Pires, Bernardo Ávila, et al.
Publicado: (2025)
por: Pires, Bernardo Ávila, et al.
Publicado: (2025)
The Return of Pseudosciences in Artificial Intelligence: Have Machine Learning and Deep Learning Forgotten Lessons from Statistics and History?
por: Sublime, Jérémie
Publicado: (2024)
por: Sublime, Jérémie
Publicado: (2024)
Reinforcement Learning-Guided Semi-Supervised Learning
por: Heidari, Marzi, et al.
Publicado: (2024)
por: Heidari, Marzi, et al.
Publicado: (2024)
Chunk-Guided Q-Learning
por: Song, Gwanwoo, et al.
Publicado: (2026)
por: Song, Gwanwoo, et al.
Publicado: (2026)
Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling
por: Sinha, Abhijeet, et al.
Publicado: (2026)
por: Sinha, Abhijeet, et al.
Publicado: (2026)
D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks
por: Zhang, Jundong, et al.
Publicado: (2025)
por: Zhang, Jundong, et al.
Publicado: (2025)
RGMDT: Return-Gap-Minimizing Decision Tree Extraction in Non-Euclidean Metric Space
por: Chen, Jingdi, et al.
Publicado: (2024)
por: Chen, Jingdi, et al.
Publicado: (2024)
R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning
por: Farhi, Nadir
Publicado: (2025)
por: Farhi, Nadir
Publicado: (2025)
Interpretable Deep Learning for Stock Returns: A Consensus-Bottleneck Asset Pricing Model
por: Kim, Changeun, et al.
Publicado: (2025)
por: Kim, Changeun, et al.
Publicado: (2025)
Provable and Practical In-Context Policy Optimization for Self-Improvement
por: Yu, Tianrun, et al.
Publicado: (2026)
por: Yu, Tianrun, et al.
Publicado: (2026)
Labels Matter More Than Models: Rethinking the Unsupervised Paradigm in Time Series Anomaly Detection
por: Zhong, Zhijie, et al.
Publicado: (2025)
por: Zhong, Zhijie, et al.
Publicado: (2025)
LCDB 1.1: A Database Illustrating Learning Curves Are More Ill-Behaved Than Previously Thought
por: Yan, Cheng, et al.
Publicado: (2025)
por: Yan, Cheng, et al.
Publicado: (2025)
Joint Return and Risk Modeling with Deep Neural Networks for Portfolio Construction
por: Park, Keonvin
Publicado: (2026)
por: Park, Keonvin
Publicado: (2026)
Decoupling Return-to-Go for Efficient Decision Transformer
por: Wang, Yongyi, et al.
Publicado: (2026)
por: Wang, Yongyi, et al.
Publicado: (2026)
Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration
por: Zhang, Ziqi, et al.
Publicado: (2023)
por: Zhang, Ziqi, et al.
Publicado: (2023)
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
por: Li, Weiqi, et al.
Publicado: (2025)
por: Li, Weiqi, et al.
Publicado: (2025)
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
por: Zheng, Liangwei Nathan, et al.
Publicado: (2025)
por: Zheng, Liangwei Nathan, et al.
Publicado: (2025)
Why DDIM Hallucinates More Than DDPM: A Theoretical Analysis of Reverse Dynamics
por: Ashiq, Muhammad H., et al.
Publicado: (2026)
por: Ashiq, Muhammad H., et al.
Publicado: (2026)
Learning to Spend: Model Predictive Control for Budgeting under Non-Stationary Returns
por: Pathak, Nilavra, et al.
Publicado: (2026)
por: Pathak, Nilavra, et al.
Publicado: (2026)
Explainable AI for Mental Health Emergency Returns: Integrating LLMs with Predictive Modeling
por: Ahmed, Abdulaziz, et al.
Publicado: (2025)
por: Ahmed, Abdulaziz, et al.
Publicado: (2025)
Curriculum Is More Influential Than Haptic Information During Reinforcement Learning of Object Manipulation Against Gravity
por: Ojaghi, Pegah, et al.
Publicado: (2024)
por: Ojaghi, Pegah, et al.
Publicado: (2024)
Ejemplares similares
-
How to Provably Improve Return Conditioned Supervised Learning?
por: Liu, Zhishuai, et al.
Publicado: (2025) -
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
por: Yang, Yuxiao, et al.
Publicado: (2026) -
Imitation Learning as Return Distribution Matching
por: Lazzati, Filippo, et al.
Publicado: (2025) -
On the Diminishing Returns of Width for Continual Learning
por: Guha, Etash, et al.
Publicado: (2024) -
STAS: Spatial-Temporal Return Decomposition for Multi-agent Reinforcement Learning
por: Chen, Sirui, et al.
Publicado: (2023)