Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Fan, Jia, Zeyu, Rakhlin, Alexander, Xie, Tengyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
Near-Optimal Learning and Planning in Separated Latent MDPs
by: Chen, Fan, et al.
Published: (2024)
by: Chen, Fan, et al.
Published: (2024)
Beyond Covariance Matrix: The Statistical Complexity of Private Linear Regression
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
by: Chen, Fan, et al.
Published: (2022)
by: Chen, Fan, et al.
Published: (2022)
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025)
by: Chi, Yuejie, et al.
Published: (2025)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Online Learning with Unknown Constraints
by: Sridharan, Karthik, et al.
Published: (2024)
by: Sridharan, Karthik, et al.
Published: (2024)
The Power of Resets in Online Reinforcement Learning
by: Mhammedi, Zakaria, et al.
Published: (2024)
by: Mhammedi, Zakaria, et al.
Published: (2024)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Self-Normalized Martingales and Uniform Regret Bounds for Linear Regression
by: Chen, Fan, et al.
Published: (2026)
by: Chen, Fan, et al.
Published: (2026)
High-accuracy sampling for diffusion models and log-concave distributions
by: Chen, Fan, et al.
Published: (2026)
by: Chen, Fan, et al.
Published: (2026)
Online Estimation via Offline Estimation: An Information-Theoretic Framework
by: Foster, Dylan J., et al.
Published: (2024)
by: Foster, Dylan J., et al.
Published: (2024)
Conformal Prediction for Privacy-Preserving Machine Learning
by: Balinsky, Alexander David, et al.
Published: (2025)
by: Balinsky, Alexander David, et al.
Published: (2025)
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
by: Bian, Zeyu, et al.
Published: (2026)
by: Bian, Zeyu, et al.
Published: (2026)
Scaling Limits of Long-Context Transformers
by: Bruno, Giuseppe, et al.
Published: (2026)
by: Bruno, Giuseppe, et al.
Published: (2026)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
by: Shen, Zhaiming, et al.
Published: (2025)
by: Shen, Zhaiming, et al.
Published: (2025)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
by: Baharav, Tavor Z., et al.
Published: (2025)
by: Baharav, Tavor Z., et al.
Published: (2025)
On the Variance, Admissibility, and Stability of Empirical Risk Minimization
by: Kur, Gil, et al.
Published: (2023)
by: Kur, Gil, et al.
Published: (2023)
Refined Risk Bounds for Unbounded Losses via Transductive Priors
by: Qian, Jian, et al.
Published: (2024)
by: Qian, Jian, et al.
Published: (2024)
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
by: Foster, Dylan J., et al.
Published: (2025)
by: Foster, Dylan J., et al.
Published: (2025)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning
by: Cao, Junyu, et al.
Published: (2026)
by: Cao, Junyu, et al.
Published: (2026)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
by: Lu, Miao, et al.
Published: (2022)
by: Lu, Miao, et al.
Published: (2022)
A Differential and Pointwise Control Approach to Reinforcement Learning
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026)
by: Nguyen, Minh
Published: (2026)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
by: Rajendran, Goutham, et al.
Published: (2024)
by: Rajendran, Goutham, et al.
Published: (2024)
Compression, Generalization and Learning
by: Campi, Marco C., et al.
Published: (2023)
by: Campi, Marco C., et al.
Published: (2023)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Adaptive Sample Aggregation In Transfer Learning
by: Hanneke, Steve, et al.
Published: (2024)
by: Hanneke, Steve, et al.
Published: (2024)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
Learning with Differentially Private (Sliced) Wasserstein Gradients
by: Rodríguez-Vítores, David, et al.
Published: (2025)
by: Rodríguez-Vítores, David, et al.
Published: (2025)
Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability
by: Yu, Lijia, et al.
Published: (2025)
by: Yu, Lijia, et al.
Published: (2025)
Similar Items
-
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
by: Chen, Fan, et al.
Published: (2025) -
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
by: Jia, Zeyu, et al.
Published: (2025) -
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025) -
Near-Optimal Learning and Planning in Separated Latent MDPs
by: Chen, Fan, et al.
Published: (2024) -
Beyond Covariance Matrix: The Statistical Complexity of Private Linear Regression
by: Chen, Fan, et al.
Published: (2025)