Saved in:
| Main Authors: | Jia, Zeyu, Rakhlin, Alexander, Sekhari, Ayush, Wei, Chen-Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.17091 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GaussMark: A Practical Approach for Structural Watermarking of Language Models
by: Block, Adam, et al.
Published: (2025)
by: Block, Adam, et al.
Published: (2025)
The Role of Environment Access in Agnostic Reinforcement Learning
by: Krishnamurthy, Akshay, et al.
Published: (2025)
by: Krishnamurthy, Akshay, et al.
Published: (2025)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
Random Latent Exploration for Deep Reinforcement Learning
by: Mahankali, Srinath, et al.
Published: (2024)
by: Mahankali, Srinath, et al.
Published: (2024)
Offline Trajectory Optimization for Offline Reinforcement Learning
by: Zhao, Ziqi, et al.
Published: (2024)
by: Zhao, Ziqi, et al.
Published: (2024)
The Power of Resets in Online Reinforcement Learning
by: Mhammedi, Zakaria, et al.
Published: (2024)
by: Mhammedi, Zakaria, et al.
Published: (2024)
Offline Reinforcement Learning with Generative Trajectory Policies
by: Feng, Xinsong, et al.
Published: (2025)
by: Feng, Xinsong, et al.
Published: (2025)
In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning
by: Tu, Songjun, et al.
Published: (2024)
by: Tu, Songjun, et al.
Published: (2024)
When Less is Enough: Efficient Inference via Collaborative Reasoning
by: Chen, Yilei, et al.
Published: (2026)
by: Chen, Yilei, et al.
Published: (2026)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Offline Safe Reinforcement Learning Using Trajectory Classification
by: Gong, Ze, et al.
Published: (2024)
by: Gong, Ze, et al.
Published: (2024)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
by: Hu, Hao, et al.
Published: (2025)
by: Hu, Hao, et al.
Published: (2025)
TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning
by: Sestini, Alessandro, et al.
Published: (2025)
by: Sestini, Alessandro, et al.
Published: (2025)
State-Constrained Offline Reinforcement Learning
by: Hepburn, Charles A., et al.
Published: (2024)
by: Hepburn, Charles A., et al.
Published: (2024)
GTA: Generative Trajectory Augmentation with Guidance for Offline Reinforcement Learning
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
by: Fang, Zeyu, et al.
Published: (2026)
by: Fang, Zeyu, et al.
Published: (2026)
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage
by: Liu, Haolin, et al.
Published: (2026)
by: Liu, Haolin, et al.
Published: (2026)
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
by: Huang, Kevin, et al.
Published: (2025)
by: Huang, Kevin, et al.
Published: (2025)
Pessimistic Causal Reinforcement Learning with Mediators for Confounded Offline Data
by: Wang, Danyang, et al.
Published: (2024)
by: Wang, Danyang, et al.
Published: (2024)
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
by: Mao, Yixiu, et al.
Published: (2024)
by: Mao, Yixiu, et al.
Published: (2024)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
by: Madhow, Sunil, et al.
Published: (2023)
by: Madhow, Sunil, et al.
Published: (2023)
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
by: Li, Guanghe, et al.
Published: (2024)
by: Li, Guanghe, et al.
Published: (2024)
Permutation Equivariant Model-based Offline Reinforcement Learning for Auto-bidding
by: Mou, Zhiyu, et al.
Published: (2025)
by: Mou, Zhiyu, et al.
Published: (2025)
Abstraction for Offline Goal-Conditioned Reinforcement Learning
by: Wibault, Clarisse, et al.
Published: (2026)
by: Wibault, Clarisse, et al.
Published: (2026)
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2024)
by: Gao, Chen-Xiao, et al.
Published: (2024)
Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement Learning
by: Wen, Xiaoyu, et al.
Published: (2024)
by: Wen, Xiaoyu, et al.
Published: (2024)
SAMG: Offline-to-Online Reinforcement Learning via State-Action-Conditional Offline Model Guidance
by: Zhang, Liyu, et al.
Published: (2024)
by: Zhang, Liyu, et al.
Published: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
Consistency Trajectory Planning: High-Quality and Efficient Trajectory Optimization for Offline Model-Based Reinforcement Learning
by: Wang, Guanquan, et al.
Published: (2025)
by: Wang, Guanquan, et al.
Published: (2025)
Variational OOD State Correction for Offline Reinforcement Learning
by: Jiang, Ke, et al.
Published: (2025)
by: Jiang, Ke, et al.
Published: (2025)
Mildly Conservative Regularized Evaluation for Offline Reinforcement Learning
by: Chen, Haohui, et al.
Published: (2025)
by: Chen, Haohui, et al.
Published: (2025)
Near-Optimal Learning and Planning in Separated Latent MDPs
by: Chen, Fan, et al.
Published: (2024)
by: Chen, Fan, et al.
Published: (2024)
Pre-training with Synthetic Data Helps Offline Reinforcement Learning
by: Wang, Zecheng, et al.
Published: (2023)
by: Wang, Zecheng, et al.
Published: (2023)
FOVA: Offline Federated Reinforcement Learning with Mixed-Quality Data
by: Qiao, Nan, et al.
Published: (2025)
by: Qiao, Nan, et al.
Published: (2025)
Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
by: Golowich, Noah, et al.
Published: (2024)
by: Golowich, Noah, et al.
Published: (2024)
Similar Items
-
GaussMark: A Practical Approach for Structural Watermarking of Language Models
by: Block, Adam, et al.
Published: (2025) -
The Role of Environment Access in Agnostic Reinforcement Learning
by: Krishnamurthy, Akshay, et al.
Published: (2025) -
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
by: Chen, Fan, et al.
Published: (2025) -
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
by: Jia, Zeyu, et al.
Published: (2025) -
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025)