Saved in:
| Main Authors: | Chen, Yatong, Tang, Wei, Ho, Chien-Ju, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2305.01094 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Early Prediction of Future Behavioral Strategy from Process Traces
by: Kasumba, Robert, et al.
Published: (2026)
by: Kasumba, Robert, et al.
Published: (2026)
Bandit and Delayed Feedback in Online Structured Prediction
by: Shibukawa, Yuki, et al.
Published: (2025)
by: Shibukawa, Yuki, et al.
Published: (2025)
Reparameterization through Coverings and Topological Weight Priors
by: Beketov, Maxim, et al.
Published: (2026)
by: Beketov, Maxim, et al.
Published: (2026)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
by: Wei, Wang, et al.
Published: (2025)
by: Wei, Wang, et al.
Published: (2025)
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
by: Ge, Haosen, et al.
Published: (2024)
by: Ge, Haosen, et al.
Published: (2024)
Learning Intractable Multimodal Policies with Reparameterization and Diversity Regularization
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
by: Zhao, Boxin, et al.
Published: (2021)
by: Zhao, Boxin, et al.
Published: (2021)
Learning Equilibria in Matching Games with Bandit Feedback
by: Athanasopoulos, Andreas, et al.
Published: (2025)
by: Athanasopoulos, Andreas, et al.
Published: (2025)
Optimal Clustering with Bandit Feedback
by: Yang, Junwen, et al.
Published: (2022)
by: Yang, Junwen, et al.
Published: (2022)
On Mitigating Affinity Bias through Bandits with Evolving Biased Feedback
by: Faw, Matthew, et al.
Published: (2025)
by: Faw, Matthew, et al.
Published: (2025)
Lipschitz Bandits with Stochastic Delayed Feedback
by: Liu, Zhongxuan, et al.
Published: (2025)
by: Liu, Zhongxuan, et al.
Published: (2025)
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
Learning to Schedule Online Tasks with Bandit Feedback
by: Xu, Yongxin, et al.
Published: (2024)
by: Xu, Yongxin, et al.
Published: (2024)
Does Feedback Help in Bandits with Arm Erasures?
by: Karakas, Merve, et al.
Published: (2025)
by: Karakas, Merve, et al.
Published: (2025)
Nearest Neighbour with Bandit Feedback
by: Pasteris, Stephen, et al.
Published: (2023)
by: Pasteris, Stephen, et al.
Published: (2023)
Strategic Hypothesis Testing
by: Hossain, Safwan, et al.
Published: (2025)
by: Hossain, Safwan, et al.
Published: (2025)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
by: Goyal, Tanmay, et al.
Published: (2025)
by: Goyal, Tanmay, et al.
Published: (2025)
Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting
by: Yang, Sifan, et al.
Published: (2025)
by: Yang, Sifan, et al.
Published: (2025)
Linear Submodular Maximization with Bandit Feedback
by: Chen, Wenjing, et al.
Published: (2024)
by: Chen, Wenjing, et al.
Published: (2024)
Clustering Items through Bandit Feedback: Finding the Right Feature out of Many
by: Graf, Maximilian, et al.
Published: (2025)
by: Graf, Maximilian, et al.
Published: (2025)
Learning Markov Decision Processes under Fully Bandit Feedback
by: Zhuo, Zhengjia, et al.
Published: (2026)
by: Zhuo, Zhengjia, et al.
Published: (2026)
Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits
by: Li, Shaoang, et al.
Published: (2025)
by: Li, Shaoang, et al.
Published: (2025)
Queueing Matching Bandits with Preference Feedback
by: Kim, Jung-hun, et al.
Published: (2024)
by: Kim, Jung-hun, et al.
Published: (2024)
Graph Feedback Bandits with Similar Arms
by: Qi, Han, et al.
Published: (2024)
by: Qi, Han, et al.
Published: (2024)
Nonparametric Kernel Clustering with Bandit Feedback
by: Thuot, Victor, et al.
Published: (2026)
by: Thuot, Victor, et al.
Published: (2026)
Cascading Bandits With Feedback
by: Prakash, R Sri, et al.
Published: (2025)
by: Prakash, R Sri, et al.
Published: (2025)
To Give or Not to Give? The Impacts of Strategically Withheld Recourse
by: Chen, Yatong, et al.
Published: (2025)
by: Chen, Yatong, et al.
Published: (2025)
ReDiSC: A Reparameterized Masked Diffusion Model for Scalable Node Classification with Structured Predictions
by: Li, Yule, et al.
Published: (2025)
by: Li, Yule, et al.
Published: (2025)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
by: Truong, Tuan, et al.
Published: (2025)
by: Truong, Tuan, et al.
Published: (2025)
Stochastic $k$-Submodular Bandits with Full Bandit Feedback
by: Nie, Guanyu, et al.
Published: (2024)
by: Nie, Guanyu, et al.
Published: (2024)
Non-Stationary Bandit Learning via Predictive Sampling
by: Liu, Yueyang, et al.
Published: (2022)
by: Liu, Yueyang, et al.
Published: (2022)
Categorical Reparameterization with Denoising Diffusion models
by: Gourevitch, Samson, et al.
Published: (2026)
by: Gourevitch, Samson, et al.
Published: (2026)
Reparameterization invariance in approximate Bayesian inference
by: Roy, Hrittik, et al.
Published: (2024)
by: Roy, Hrittik, et al.
Published: (2024)
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
by: Li, Zitian, et al.
Published: (2026)
by: Li, Zitian, et al.
Published: (2026)
Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback
by: Huang, Ruiyuan, et al.
Published: (2025)
by: Huang, Ruiyuan, et al.
Published: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025)
by: Zhong, Hai, et al.
Published: (2025)
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026)
by: Zhong, Hai, et al.
Published: (2026)
Multiclass Online Learnability under Bandit Feedback
by: Raman, Ananth, et al.
Published: (2023)
by: Raman, Ananth, et al.
Published: (2023)
Similar Items
-
Early Prediction of Future Behavioral Strategy from Process Traces
by: Kasumba, Robert, et al.
Published: (2026) -
Bandit and Delayed Feedback in Online Structured Prediction
by: Shibukawa, Yuki, et al.
Published: (2025) -
Reparameterization through Coverings and Topological Weight Priors
by: Beketov, Maxim, et al.
Published: (2026) -
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024) -
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
by: Wei, Wang, et al.
Published: (2025)