Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ayoub, Alex, Wang, Kaiwen, Liu, Vincent, Robertson, Samuel, McInerney, James, Liang, Dawen, Kallus, Nathan, Szepesvári, Csaba |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
por: McInerney, James, et al.
Publicado: (2024)
por: McInerney, James, et al.
Publicado: (2024)
Does Weighting Improve Matrix Factorization for Recommender Systems?
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
por: Wang, Kaiwen, et al.
Publicado: (2024)
por: Wang, Kaiwen, et al.
Publicado: (2024)
Rectifying Regression in Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Entropy After </Think> for reasoning model early exiting
por: Wang, Xi, et al.
Publicado: (2025)
por: Wang, Xi, et al.
Publicado: (2025)
Optimization of Epsilon-Greedy Exploration
por: Che, Ethan, et al.
Publicado: (2025)
por: Che, Ethan, et al.
Publicado: (2025)
Adjusting Regression Models for Conditional Uncertainty Calibration
por: Gao, Ruijiang, et al.
Publicado: (2024)
por: Gao, Ruijiang, et al.
Publicado: (2024)
The Central Role of the Loss Function in Reinforcement Learning
por: Wang, Kaiwen, et al.
Publicado: (2024)
por: Wang, Kaiwen, et al.
Publicado: (2024)
Eluder dimension: localise it!
por: Bakhtiari, Alireza, et al.
Publicado: (2026)
por: Bakhtiari, Alireza, et al.
Publicado: (2026)
Learning to Reason Efficiently with Discounted Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Exploration via linearly perturbed loss minimisation
por: Janz, David, et al.
Publicado: (2023)
por: Janz, David, et al.
Publicado: (2023)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
por: Liu, Shuai, et al.
Publicado: (2026)
por: Liu, Shuai, et al.
Publicado: (2026)
Teaching Knowledge Management (SIG KM).
por: McInerney, Claire
Publicado: (2000)
por: McInerney, Claire
Publicado: (2000)
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
por: Zhao, Hanyang, et al.
Publicado: (2025)
por: Zhao, Hanyang, et al.
Publicado: (2025)
Bellman Calibration for $V$-Learning in Offline Reinforcement Learning
por: van der Laan, Lars, et al.
Publicado: (2025)
por: van der Laan, Lars, et al.
Publicado: (2025)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
por: Tkachuk, Volodymyr, et al.
Publicado: (2024)
por: Tkachuk, Volodymyr, et al.
Publicado: (2024)
Environmental Scanning and the Information Manager.
por: Newsome, James, et al.
Publicado: (1990)
por: Newsome, James, et al.
Publicado: (1990)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
por: Liu, Shuai, et al.
Publicado: (2024)
por: Liu, Shuai, et al.
Publicado: (2024)
Reproductive behaviour of the blackspotted stickleback, Gasterostomus wheatlandi
por: McInerney, J. E
Publicado: (1969)
por: McInerney, J. E
Publicado: (1969)
Justice, Complexity and Effective Governance in the Twenty-First Century
por: Thomas F. McInerney
Publicado: (2021)
por: Thomas F. McInerney
Publicado: (2021)
A Statistical-Modelling Approach to Feedforward Neural Network Model Selection
por: McInerney, Andrew, et al.
Publicado: (2022)
por: McInerney, Andrew, et al.
Publicado: (2022)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
por: Kitamura, Toshinori, et al.
Publicado: (2026)
por: Kitamura, Toshinori, et al.
Publicado: (2026)
Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
por: Zhu, Yaochen, et al.
Publicado: (2025)
por: Zhu, Yaochen, et al.
Publicado: (2025)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
por: Maran, Davide, et al.
Publicado: (2026)
por: Maran, Davide, et al.
Publicado: (2026)
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
por: Wang, Kaiwen, et al.
Publicado: (2024)
por: Wang, Kaiwen, et al.
Publicado: (2024)
How Has the Rise of Direct‐To‐Consumer Genetic Testing Impacted Genetic Counselling Practice? A Scoping Review
por: Cushla McKinney, et al.
Publicado: (2026)
por: Cushla McKinney, et al.
Publicado: (2026)
A High-level Synthesis Toolchain for the Julia Language
por: Short, Benedict, et al.
Publicado: (2025)
por: Short, Benedict, et al.
Publicado: (2025)
Hardware.jl - An MLIR-based Julia HLS Flow (Work in Progress)
por: Short, Benedict, et al.
Publicado: (2025)
por: Short, Benedict, et al.
Publicado: (2025)
Action spectrum of the photoperiod mechanism controlling sexual maturation in the threespine stickleback, Gasterosteus acuelatus
por: McInerney, J. E. and Evans, D. O
Publicado: (1970)
por: McInerney, J. E. and Evans, D. O
Publicado: (1970)
Habitat characteristics of the Pacific hagfish, Polistotrema stouti
por: McInerney, J. E., Evans, D. O
Publicado: (1970)
por: McInerney, J. E., Evans, D. O
Publicado: (1970)
Development of salinity preference in pre-smolt coho salmon, Oncorhynchus kisutch
por: Otto, R. G., McInerney, J. E
Publicado: (1970)
por: Otto, R. G., McInerney, J. E
Publicado: (1970)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
por: Liu, Chang, et al.
Publicado: (2026)
por: Liu, Chang, et al.
Publicado: (2026)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
por: György, András, et al.
Publicado: (2025)
por: György, András, et al.
Publicado: (2025)
Balancing optimism and pessimism in offline-to-online learning
por: Sentenac, Flore, et al.
Publicado: (2025)
por: Sentenac, Flore, et al.
Publicado: (2025)
Sharp analysis of linear ensemble sampling
por: Akhavan, Arya, et al.
Publicado: (2026)
por: Akhavan, Arya, et al.
Publicado: (2026)
Reindex-Then-Adapt: Improving Large Language Models for Conversational Recommendation
por: He, Zhankui, et al.
Publicado: (2024)
por: He, Zhankui, et al.
Publicado: (2024)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
por: Kallus, Nathan
Publicado: (2025)
por: Kallus, Nathan
Publicado: (2025)
To Switch or Not to Switch? Balanced Policy Switching in Offline Reinforcement Learning
por: Ma, Tao, et al.
Publicado: (2024)
por: Ma, Tao, et al.
Publicado: (2024)
Collaborative Retrieval for Large Language Model-based Conversational Recommender Systems
por: Zhu, Yaochen, et al.
Publicado: (2025)
por: Zhu, Yaochen, et al.
Publicado: (2025)
Chapter New Lenses
por: McInerney, William W., et al.
Publicado: (2023)
por: McInerney, William W., et al.
Publicado: (2023)
Ejemplares similares
-
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
por: McInerney, James, et al.
Publicado: (2024) -
Does Weighting Improve Matrix Factorization for Recommender Systems?
por: Ayoub, Alex, et al.
Publicado: (2025) -
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
por: Wang, Kaiwen, et al.
Publicado: (2024) -
Rectifying Regression in Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025) -
Entropy After </Think> for reasoning model early exiting
por: Wang, Xi, et al.
Publicado: (2025)