Multi-Player Approaches for Dueling Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Raveh, Or, Honda, Junya, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Survival Bandit Problem
by: Riou, Charles, et al.
Published: (2022)
by: Riou, Charles, et al.
Published: (2022)
Thompson Exploration with Best Challenger Rule in Best Arm Identification
by: Lee, Jongyeong, et al.
Published: (2023)
by: Lee, Jongyeong, et al.
Published: (2023)
LLM Routing with Dueling Feedback
by: Chiang, Chao-Kai, et al.
Published: (2025)
by: Chiang, Chao-Kai, et al.
Published: (2025)
A Fast Algorithm for the Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit
by: Nakamura, Shintaro, et al.
Published: (2023)
by: Nakamura, Shintaro, et al.
Published: (2023)
Note on Follow-the-Perturbed-Leader in Combinatorial Semi-Bandit Problems
by: Chen, Botao, et al.
Published: (2025)
by: Chen, Botao, et al.
Published: (2025)
A General Recipe for the Analysis of Randomized Multi-Armed Bandit Algorithms
by: Baudry, Dorian, et al.
Published: (2023)
by: Baudry, Dorian, et al.
Published: (2023)
Federated Linear Dueling Bandits
by: Huang, Xuhan, et al.
Published: (2025)
by: Huang, Xuhan, et al.
Published: (2025)
Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare
by: Ahmed, Maheed H., et al.
Published: (2026)
by: Ahmed, Maheed H., et al.
Published: (2026)
Online Clustering of Dueling Bandits
by: Wang, Zhiyong, et al.
Published: (2025)
by: Wang, Zhiyong, et al.
Published: (2025)
Biased Dueling Bandits with Stochastic Delayed Feedback
by: Yi, Bongsoo, et al.
Published: (2024)
by: Yi, Bongsoo, et al.
Published: (2024)
Revisiting Follow-the-Perturbed-Leader with Unbounded Perturbations in Bandit Problems
by: Lee, Jongyeong, et al.
Published: (2025)
by: Lee, Jongyeong, et al.
Published: (2025)
A Further Efficient Algorithm with Best-of-Both-Worlds Guarantees for $m$-Set Semi-Bandit Problem
by: Chen, Botao, et al.
Published: (2026)
by: Chen, Botao, et al.
Published: (2026)
The Sampling Complexity of Condorcet Winner Identification in Dueling Bandits
by: Saad, El Mehdi, et al.
Published: (2026)
by: Saad, El Mehdi, et al.
Published: (2026)
Conversational Dueling Bandits in Generalized Linear Models
by: Yang, Shuhua, et al.
Published: (2024)
by: Yang, Shuhua, et al.
Published: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)
by: Wang, Xiangyi, et al.
Published: (2026)
Follow-the-Perturbed-Leader with Fréchet-type Tail Distributions: Optimality in Adversarial Bandits and Best-of-Both-Worlds
by: Lee, Jongyeong, et al.
Published: (2024)
by: Lee, Jongyeong, et al.
Published: (2024)
Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
by: Zhang, Yu-Jie, et al.
Published: (2025)
by: Zhang, Yu-Jie, et al.
Published: (2025)
Utility-based Dueling Bandits as a Partial Monitoring Game
by: Gajane, Pratik, et al.
Published: (2015)
by: Gajane, Pratik, et al.
Published: (2015)
Recycling History: Efficient Recommendations from Contextual Dueling Bandits
by: Sankagiri, Suryanarayana, et al.
Published: (2025)
by: Sankagiri, Suryanarayana, et al.
Published: (2025)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024)
by: Li, Xuheng, et al.
Published: (2024)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
Non-Stationary Dueling Bandits Under a Weighted Borda Criterion
by: Suk, Joe, et al.
Published: (2024)
by: Suk, Joe, et al.
Published: (2024)
Active Human Feedback Collection via Neural Contextual Dueling Bandits
by: Verma, Arun, et al.
Published: (2025)
by: Verma, Arun, et al.
Published: (2025)
When Can We Track Significant Preference Shifts in Dueling Bandits?
by: Suk, Joe, et al.
Published: (2023)
by: Suk, Joe, et al.
Published: (2023)
Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration
by: Oh, Youngmin, et al.
Published: (2025)
by: Oh, Youngmin, et al.
Published: (2025)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Lipschitz Dueling Bandits over Continuous Action Spaces
by: Sharma, Mudit, et al.
Published: (2026)
by: Sharma, Mudit, et al.
Published: (2026)
Expected Possession Value of Control and Duel Actions for Soccer Player's Skills Estimation
by: Shelopugin, Andrei
Published: (2024)
by: Shelopugin, Andrei
Published: (2024)
Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks
by: Magesh, Akshayaa, et al.
Published: (2025)
by: Magesh, Akshayaa, et al.
Published: (2025)
Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
Best-of-Both-Worlds Multi-Dueling Bandits: Unified Algorithms for Stochastic and Adversarial Preferences under Condorcet and Borda Objectives
by: Akash, S, et al.
Published: (2026)
by: Akash, S, et al.
Published: (2026)
Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions
by: Oh, Youngmin
Published: (2026)
by: Oh, Youngmin
Published: (2026)
Exploration by Optimization with Hybrid Regularizers: Logarithmic Regret with Adversarial Robustness in Partial Monitoring
by: Tsuchiya, Taira, et al.
Published: (2024)
by: Tsuchiya, Taira, et al.
Published: (2024)
Adaptive Learning Rate for Follow-the-Regularized-Leader: Competitive Analysis and Best-of-Both-Worlds
by: Ito, Shinji, et al.
Published: (2024)
by: Ito, Shinji, et al.
Published: (2024)
Rate-optimal Design for Anytime Best Arm Identification
by: Komiyama, Junpei, et al.
Published: (2025)
by: Komiyama, Junpei, et al.
Published: (2025)
Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds
by: Tsuchiya, Taira, et al.
Published: (2023)
by: Tsuchiya, Taira, et al.
Published: (2023)
Optimal Regret of Bernoulli Bandits under Global Differential Privacy
by: Azize, Achraf, et al.
Published: (2025)
by: Azize, Achraf, et al.
Published: (2025)
Similar Items
-
The Survival Bandit Problem
by: Riou, Charles, et al.
Published: (2022) -
Thompson Exploration with Best Challenger Rule in Best Arm Identification
by: Lee, Jongyeong, et al.
Published: (2023) -
LLM Routing with Dueling Feedback
by: Chiang, Chao-Kai, et al.
Published: (2025) -
A Fast Algorithm for the Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit
by: Nakamura, Shintaro, et al.
Published: (2023) -
Note on Follow-the-Perturbed-Leader in Combinatorial Semi-Bandit Problems
by: Chen, Botao, et al.
Published: (2025)