Fusing Reward and Dueling Feedback in Stochastic Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xuchuang, Zeng, Qirun, Zuo, Jinhang, Liu, Xutong, Hajiesmaili, Mohammad, Lui, John C. S., Wierman, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023)
by: Liu, Xutong, et al.
Published: (2023)
Stochastic Bandits Robust to Adversarial Attacks
by: Wang, Xuchuang, et al.
Published: (2024)
by: Wang, Xuchuang, et al.
Published: (2024)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
by: Zeng, Qirun, et al.
Published: (2025)
by: Zeng, Qirun, et al.
Published: (2025)
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
by: Zeng, Qirun, et al.
Published: (2026)
by: Zeng, Qirun, et al.
Published: (2026)
Multi-Agent Stochastic Bandits Robust to Adversarial Corruptions
by: Ghaffari, Fatemeh, et al.
Published: (2024)
by: Ghaffari, Fatemeh, et al.
Published: (2024)
Heterogeneous Multi-Agent Bandits with Parsimonious Hints
by: Mirfakhar, Amirmahdi, et al.
Published: (2025)
by: Mirfakhar, Amirmahdi, et al.
Published: (2025)
Combinatorial Logistic Bandits
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
by: Liu, Xutong, et al.
Published: (2022)
by: Liu, Xutong, et al.
Published: (2022)
Online Clustering of Dueling Bandits
by: Wang, Zhiyong, et al.
Published: (2025)
by: Wang, Zhiyong, et al.
Published: (2025)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
Unlearning Offline Stochastic Multi-Armed Bandits
by: Ye, Zichun, et al.
Published: (2026)
by: Ye, Zichun, et al.
Published: (2026)
Heterogeneous Multi-agent Multi-armed Bandits on Stochastic Block Models
by: Xu, Mengfan, et al.
Published: (2025)
by: Xu, Mengfan, et al.
Published: (2025)
Offline Clustering of Linear Bandits: The Power of Clusters under Limited Data
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)
by: Wang, Xiangyi, et al.
Published: (2026)
Online Learning to Rank under Corruption: A Robust Cascading Bandits Approach
by: Ghaffari, Fatemeh, et al.
Published: (2025)
by: Ghaffari, Fatemeh, et al.
Published: (2025)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Competitive Algorithms for Multi-Agent Ski-Rental Problems
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
Towards Environmentally Equitable AI
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
by: Yang, Hantao, et al.
Published: (2024)
by: Yang, Hantao, et al.
Published: (2024)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
by: Dai, Xiangxiang, et al.
Published: (2024)
by: Dai, Xiangxiang, et al.
Published: (2024)
Demystifying Online Clustering of Bandits: Enhanced Exploration Under Stochastic and Smoothed Adversarial Contexts
by: Li, Zhuohua, et al.
Published: (2025)
by: Li, Zhuohua, et al.
Published: (2025)
Offline Learning for Combinatorial Multi-armed Bandits
by: Liu, Xutong, et al.
Published: (2025)
by: Liu, Xutong, et al.
Published: (2025)
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
by: Poon, Manhin, et al.
Published: (2025)
by: Poon, Manhin, et al.
Published: (2025)
Learning Best Paths in Quantum Networks
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
by: Pedramfar, Mohammad, et al.
Published: (2023)
by: Pedramfar, Mohammad, et al.
Published: (2023)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
FedConPE: Efficient Federated Conversational Bandits with Heterogeneous Clients
by: Li, Zhuohua, et al.
Published: (2024)
by: Li, Zhuohua, et al.
Published: (2024)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
HiLoRA: Adaptive Hierarchical LoRA Routing for Training-Free Domain Generalization
by: Han, Ziyi, et al.
Published: (2025)
by: Han, Ziyi, et al.
Published: (2025)
Biased Dueling Bandits with Stochastic Delayed Feedback
by: Yi, Bongsoo, et al.
Published: (2024)
by: Yi, Bongsoo, et al.
Published: (2024)
Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual Bandits
by: Liu, Maoli, et al.
Published: (2025)
by: Liu, Maoli, et al.
Published: (2025)
Offline Clustering of Preference Learning with Active-data Augmentation
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Large Language Model-Enhanced Multi-Armed Bandits
by: Sun, Jiahang, et al.
Published: (2025)
by: Sun, Jiahang, et al.
Published: (2025)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
by: Lu, Xiaodong, et al.
Published: (2026)
by: Lu, Xiaodong, et al.
Published: (2026)
A Controlled Study of Double DQN and Dueling DQN Under Cross-Environment Transfer
by: Nasir, Azkaa, et al.
Published: (2026)
by: Nasir, Azkaa, et al.
Published: (2026)
Fairness-Regularized Online Optimization with Switching Costs
by: Li, Pengfei, et al.
Published: (2025)
by: Li, Pengfei, et al.
Published: (2025)
Similar Items
-
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023) -
Stochastic Bandits Robust to Adversarial Attacks
by: Wang, Xuchuang, et al.
Published: (2024) -
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
by: Zeng, Qirun, et al.
Published: (2025) -
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
by: Zeng, Qirun, et al.
Published: (2026) -
Multi-Agent Stochastic Bandits Robust to Adversarial Corruptions
by: Ghaffari, Fatemeh, et al.
Published: (2024)