Saved in:
| Main Authors: | Zhang, Yixuan, Zhu, Ruihao, Xie, Qiaomin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.02762 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
by: Zhang, Yixuan, et al.
Published: (2026)
by: Zhang, Yixuan, et al.
Published: (2026)
Wasserstein-p Central Limit Theorem Rates: From Local Dependence to Markov Chains
by: Zhang, Yixuan, et al.
Published: (2026)
by: Zhang, Yixuan, et al.
Published: (2026)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
Prelimit Coupling and Steady-State Convergence of Constant-stepsize Nonsmooth Contractive SA
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant Stepsize
by: Huo, Dongyan, et al.
Published: (2024)
by: Huo, Dongyan, et al.
Published: (2024)
Stable Offline Value Function Learning with Bisimulation-based Representations
by: Pavse, Brahma S., et al.
Published: (2024)
by: Pavse, Brahma S., et al.
Published: (2024)
A Piecewise Lyapunov Analysis of Sub-quadratic SGD: Applications to Robust and Quantile Regression
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
by: Cheung, Wang Chi, et al.
Published: (2024)
by: Cheung, Wang Chi, et al.
Published: (2024)
Simulating Biases for Interpretable Fairness in Offline and Online Classifiers
by: Inácio, Ricardo, et al.
Published: (2025)
by: Inácio, Ricardo, et al.
Published: (2025)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
by: Han, Zean, et al.
Published: (2026)
by: Han, Zean, et al.
Published: (2026)
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
by: Pavse, Brahma S., et al.
Published: (2023)
by: Pavse, Brahma S., et al.
Published: (2023)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
by: Huo, Dongyan, et al.
Published: (2022)
by: Huo, Dongyan, et al.
Published: (2022)
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators
by: Linial, Ori, et al.
Published: (2024)
by: Linial, Ori, et al.
Published: (2024)
Data Poisoning to Fake a Nash Equilibrium in Markov Games
by: Wu, Young, et al.
Published: (2023)
by: Wu, Young, et al.
Published: (2023)
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
by: Chen, Ziru, et al.
Published: (2026)
by: Chen, Ziru, et al.
Published: (2026)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
Leveraging Offline Data in Linear Latent Contextual Bandits
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
Group-Sensitive Offline Contextual Bandits
by: Guo, Yihong, et al.
Published: (2025)
by: Guo, Yihong, et al.
Published: (2025)
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
C-kNN-LSH: A Nearest-Neighbor Algorithm for Sequential Counterfactual Inference
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Optimal Attack and Defense for Reinforcement Learning
by: McMahan, Jeremy, et al.
Published: (2023)
by: McMahan, Jeremy, et al.
Published: (2023)
Offline Contextual Bandit with Counterfactual Sample Identification
by: Gilotte, Alexandre, et al.
Published: (2025)
by: Gilotte, Alexandre, et al.
Published: (2025)
Offline Contextual Bandits in the Presence of New Actions
by: Kishimoto, Ren, et al.
Published: (2026)
by: Kishimoto, Ren, et al.
Published: (2026)
Online Assortment and Price Optimization Under Contextual Choice Models
by: Erginbas, Yigit Efe, et al.
Published: (2025)
by: Erginbas, Yigit Efe, et al.
Published: (2025)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Roping in Uncertainty: Robustness and Regularization in Markov Games
by: McMahan, Jeremy, et al.
Published: (2024)
by: McMahan, Jeremy, et al.
Published: (2024)
Optimally Installing Strict Equilibria
by: McMahan, Jeremy, et al.
Published: (2025)
by: McMahan, Jeremy, et al.
Published: (2025)
Contextual Dynamic Pricing with Heterogeneous Buyers
by: Lykouris, Thodoris, et al.
Published: (2025)
by: Lykouris, Thodoris, et al.
Published: (2025)
PAC-Bayes Meets Online Contextual Optimization
by: Xie, Zhuojun, et al.
Published: (2025)
by: Xie, Zhuojun, et al.
Published: (2025)
Active Advantage-Aligned Online Reinforcement Learning with Offline Data
by: Liu, Xuefeng, et al.
Published: (2025)
by: Liu, Xuefeng, et al.
Published: (2025)
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
Online Policy Learning from Offline Preferences
by: Zhang, Guoxi, et al.
Published: (2024)
by: Zhang, Guoxi, et al.
Published: (2024)
Minimax Optimality in Contextual Dynamic Pricing with General Valuation Models
by: Gong, Xueping, et al.
Published: (2024)
by: Gong, Xueping, et al.
Published: (2024)
Inception: Efficiently Computable Misinformation Attacks on Markov Games
by: McMahan, Jeremy, et al.
Published: (2024)
by: McMahan, Jeremy, et al.
Published: (2024)
Measuring and Mitigating Biases in Motor Insurance Pricing
by: Moriah, Mulah, et al.
Published: (2023)
by: Moriah, Mulah, et al.
Published: (2023)
How to Solve Contextual Goal-Oriented Problems with Offline Datasets?
by: Fan, Ying, et al.
Published: (2024)
by: Fan, Ying, et al.
Published: (2024)
Contextual Latent World Models for Offline Meta Reinforcement Learning
by: Nakheai, Mohammadreza, et al.
Published: (2026)
by: Nakheai, Mohammadreza, et al.
Published: (2026)
Online Pre-Training for Offline-to-Online Reinforcement Learning
by: Shin, Yongjae, et al.
Published: (2025)
by: Shin, Yongjae, et al.
Published: (2025)
Similar Items
-
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
by: Zhang, Yixuan, et al.
Published: (2026) -
Wasserstein-p Central Limit Theorem Rates: From Local Dependence to Markov Chains
by: Zhang, Yixuan, et al.
Published: (2026) -
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
by: Zhang, Yixuan, et al.
Published: (2024) -
Prelimit Coupling and Steady-State Convergence of Constant-stepsize Nonsmooth Contractive SA
by: Zhang, Yixuan, et al.
Published: (2024) -
The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant Stepsize
by: Huo, Dongyan, et al.
Published: (2024)