Saved in:
| Main Authors: | Zhang, Yu-Jie, Zhao, Peng, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.10616 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
by: Zhang, Yu-Jie, et al.
Published: (2025)
by: Zhang, Yu-Jie, et al.
Published: (2025)
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
by: Xie, Yan-Feng, et al.
Published: (2026)
by: Xie, Yan-Feng, et al.
Published: (2026)
Adapting to Continuous Covariate Shift via Online Density Ratio Estimation
by: Zhang, Yu-Jie, et al.
Published: (2023)
by: Zhang, Yu-Jie, et al.
Published: (2023)
Adaptivity and Non-stationarity: Problem-dependent Dynamic Regret for Online Convex Optimization
by: Zhao, Peng, et al.
Published: (2021)
by: Zhao, Peng, et al.
Published: (2021)
Mixability of Integral Losses: a Key to Efficient Online Aggregation of Functional and Probabilistic Forecasts
by: Korotin, Alexander, et al.
Published: (2019)
by: Korotin, Alexander, et al.
Published: (2019)
Efficient Methods for Non-stationary Online Learning
by: Zhao, Peng, et al.
Published: (2023)
by: Zhao, Peng, et al.
Published: (2023)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
by: Nishimori, Soichiro, et al.
Published: (2025)
by: Nishimori, Soichiro, et al.
Published: (2025)
Improved Dynamic Regret for Online Frank-Wolfe
by: Wan, Yuanyu, et al.
Published: (2023)
by: Wan, Yuanyu, et al.
Published: (2023)
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback
by: Zhang, Zhen-Yu, et al.
Published: (2026)
by: Zhang, Zhen-Yu, et al.
Published: (2026)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
by: Ackermann, Johannes, et al.
Published: (2024)
by: Ackermann, Johannes, et al.
Published: (2024)
Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
Online Structured Prediction with Fenchel--Young Losses and Improved Surrogate Regret for Online Multiclass Classification with Logistic Loss
by: Sakaue, Shinsaku, et al.
Published: (2024)
by: Sakaue, Shinsaku, et al.
Published: (2024)
Gradient-Variation Regret Bounds for Unconstrained Online Learning
by: Zhao, Yuheng, et al.
Published: (2026)
by: Zhao, Yuheng, et al.
Published: (2026)
Riemannian Langevin Dynamics: Strong Convergence of Geometric Euler-Maruyama Scheme
by: Zhan, Zhiyuan, et al.
Published: (2026)
by: Zhan, Zhiyuan, et al.
Published: (2026)
Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More Practical
by: Wang, Wei, et al.
Published: (2023)
by: Wang, Wei, et al.
Published: (2023)
Improved Approximate Regret for Decentralized Online Continuous Submodular Maximization via Reductions
by: Wan, Yuanyu, et al.
Published: (2026)
by: Wan, Yuanyu, et al.
Published: (2026)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
Enriching Disentanglement: From Logical Definitions to Quantitative Metrics
by: Zhang, Yivan, et al.
Published: (2023)
by: Zhang, Yivan, et al.
Published: (2023)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
A Fast Algorithm for the Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit
by: Nakamura, Shintaro, et al.
Published: (2023)
by: Nakamura, Shintaro, et al.
Published: (2023)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
by: Cai, Xin-Qiang, et al.
Published: (2026)
by: Cai, Xin-Qiang, et al.
Published: (2026)
Improved Kernel Alignment Regret Bound for Online Kernel Learning
by: Li, Junfan, et al.
Published: (2022)
by: Li, Junfan, et al.
Published: (2022)
Reinforcement Learning with Options and State Representation
by: Ghriss, Ayoub, et al.
Published: (2024)
by: Ghriss, Ayoub, et al.
Published: (2024)
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Improved Regret Bounds for Online Fair Division with Bandit Learning
by: Schiffer, Benjamin, et al.
Published: (2025)
by: Schiffer, Benjamin, et al.
Published: (2025)
Mixture of Online and Offline Experts for Non-stationary Time Series
by: Zhao, Zhilin, et al.
Published: (2022)
by: Zhao, Zhilin, et al.
Published: (2022)
Adaptivity and Universality: Problem-dependent Universal Regret for Online Convex Optimization
by: Zhao, Peng, et al.
Published: (2025)
by: Zhao, Peng, et al.
Published: (2025)
Achieving Optimal Static and Dynamic Regret Simultaneously in Bandits with Deterministic Losses
by: Qian, Jian, et al.
Published: (2026)
by: Qian, Jian, et al.
Published: (2026)
A Category-theoretical Meta-analysis of Definitions of Disentanglement
by: Zhang, Yivan, et al.
Published: (2023)
by: Zhang, Yivan, et al.
Published: (2023)
Multi-Player Approaches for Dueling Bandits
by: Raveh, Or, et al.
Published: (2024)
by: Raveh, Or, et al.
Published: (2024)
VEC-SBM: Optimal Community Detection with Vectorial Edges Covariates
by: Braun, Guillaume, et al.
Published: (2024)
by: Braun, Guillaume, et al.
Published: (2024)
Towards Scalable Oversight via Partitioned Human Supervision
by: Yin, Ren, et al.
Published: (2025)
by: Yin, Ren, et al.
Published: (2025)
Improved Regret in Stochastic Decision-Theoretic Online Learning under Differential Privacy
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
by: Ackermann, Johannes, et al.
Published: (2025)
by: Ackermann, Johannes, et al.
Published: (2025)
Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
by: Chen, Baiyuan, et al.
Published: (2025)
by: Chen, Baiyuan, et al.
Published: (2025)
Practical estimation of the optimal classification error with soft labels and calibration
by: Ushio, Ryota, et al.
Published: (2025)
by: Ushio, Ryota, et al.
Published: (2025)
The Survival Bandit Problem
by: Riou, Charles, et al.
Published: (2022)
by: Riou, Charles, et al.
Published: (2022)
Thompson Exploration with Best Challenger Rule in Best Arm Identification
by: Lee, Jongyeong, et al.
Published: (2023)
by: Lee, Jongyeong, et al.
Published: (2023)
Offline Reinforcement Learning with Domain-Unlabeled Data
by: Nishimori, Soichiro, et al.
Published: (2024)
by: Nishimori, Soichiro, et al.
Published: (2024)
Similar Items
-
Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
by: Zhang, Yu-Jie, et al.
Published: (2025) -
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
by: Xie, Yan-Feng, et al.
Published: (2026) -
Adapting to Continuous Covariate Shift via Online Density Ratio Estimation
by: Zhang, Yu-Jie, et al.
Published: (2023) -
Adaptivity and Non-stationarity: Problem-dependent Dynamic Regret for Online Convex Optimization
by: Zhao, Peng, et al.
Published: (2021) -
Mixability of Integral Losses: a Key to Efficient Online Aggregation of Functional and Probabilistic Forecasts
by: Korotin, Alexander, et al.
Published: (2019)