Linear and Neural Dueling Bandits with Delayed Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xiangyi, Lu, Pingchen, Mao, Jie, Kong, Mingze, Hong, Zhi, Wang, Zhiyong, Dai, Zhongxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Online Clustering of Dueling Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
von: Verma, Arun, et al.
Veröffentlicht: (2024)
von: Verma, Arun, et al.
Veröffentlicht: (2024)
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks
von: Hong, Zhi, et al.
Veröffentlicht: (2026)
von: Hong, Zhi, et al.
Veröffentlicht: (2026)
FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits
von: Lu, Pingchen, et al.
Veröffentlicht: (2025)
von: Lu, Pingchen, et al.
Veröffentlicht: (2025)
Federated Linear Dueling Bandits
von: Huang, Xuhan, et al.
Veröffentlicht: (2025)
von: Huang, Xuhan, et al.
Veröffentlicht: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
T-POP: Test-Time Personalization with Online Preference Feedback
von: Qu, Zikun, et al.
Veröffentlicht: (2025)
von: Qu, Zikun, et al.
Veröffentlicht: (2025)
Active Human Feedback Collection via Neural Contextual Dueling Bandits
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Large Language Model-Enhanced Multi-Armed Bandits
von: Sun, Jiahang, et al.
Veröffentlicht: (2025)
von: Sun, Jiahang, et al.
Veröffentlicht: (2025)
Biased Dueling Bandits with Stochastic Delayed Feedback
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
Meta-Prompt Optimization for LLM-Based Sequential Decision Making
von: Kong, Mingze, et al.
Veröffentlicht: (2025)
von: Kong, Mingze, et al.
Veröffentlicht: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
Bi-Level Contextual Bandits for Individualized Resource Allocation under Delayed Feedback
von: Almasi, Mohammadsina, et al.
Veröffentlicht: (2025)
von: Almasi, Mohammadsina, et al.
Veröffentlicht: (2025)
Conversational Dueling Bandits in Generalized Linear Models
von: Yang, Shuhua, et al.
Veröffentlicht: (2024)
von: Yang, Shuhua, et al.
Veröffentlicht: (2024)
Prompt Optimization with Human Feedback
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
Federated Linear Contextual Bandits with Heterogeneous Clients
von: Blaser, Ethan, et al.
Veröffentlicht: (2024)
von: Blaser, Ethan, et al.
Veröffentlicht: (2024)
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
von: Yang, Hantao, et al.
Veröffentlicht: (2024)
von: Yang, Hantao, et al.
Veröffentlicht: (2024)
Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
von: Flügel, Katharina, et al.
Veröffentlicht: (2023)
von: Flügel, Katharina, et al.
Veröffentlicht: (2023)
Impatient Bandits: Optimizing for the Long-Term Without Delay
von: Zhang, Kelly W., et al.
Veröffentlicht: (2025)
von: Zhang, Kelly W., et al.
Veröffentlicht: (2025)
Delayed Feedback Modeling with Influence Functions
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
Learning to Play 7 Wonders Duel Without Human Supervision
von: Paolini, Giovanni, et al.
Veröffentlicht: (2024)
von: Paolini, Giovanni, et al.
Veröffentlicht: (2024)
Conservative Contextual Bandits: Beyond Linear Representations
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
Linear Contextual Bandits with Hybrid Payoff: Revisited
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
von: Wei, Chenxing, et al.
Veröffentlicht: (2026)
von: Wei, Chenxing, et al.
Veröffentlicht: (2026)
DEER: A Delay-Resilient Framework for Reinforcement Learning with Variable Delays
von: Xia, Bo, et al.
Veröffentlicht: (2024)
von: Xia, Bo, et al.
Veröffentlicht: (2024)
Refining Adaptive Zeroth-Order Optimization at Ease
von: Shu, Yao, et al.
Veröffentlicht: (2025)
von: Shu, Yao, et al.
Veröffentlicht: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
Leveraging Offline Data in Linear Latent Contextual Bandits
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions
von: Oh, Youngmin
Veröffentlicht: (2026)
von: Oh, Youngmin
Veröffentlicht: (2026)
Neural Active Learning Beyond Bandits
von: Ban, Yikun, et al.
Veröffentlicht: (2024)
von: Ban, Yikun, et al.
Veröffentlicht: (2024)
Decentralized Online Convex Optimization with Unknown Feedback Delays
von: Qiu, Hao, et al.
Veröffentlicht: (2026)
von: Qiu, Hao, et al.
Veröffentlicht: (2026)
Learning To Play Atari Games Using Dueling Q-Learning and Hebbian Plasticity
von: Salehin, Md Ashfaq
Veröffentlicht: (2024)
von: Salehin, Md Ashfaq
Veröffentlicht: (2024)
Expected Possession Value of Control and Duel Actions for Soccer Player's Skills Estimation
von: Shelopugin, Andrei
Veröffentlicht: (2024)
von: Shelopugin, Andrei
Veröffentlicht: (2024)
Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural Networks
von: Zhao, Zhe, et al.
Veröffentlicht: (2024)
von: Zhao, Zhe, et al.
Veröffentlicht: (2024)
A Controlled Study of Double DQN and Dueling DQN Under Cross-Environment Transfer
von: Nasir, Azkaa, et al.
Veröffentlicht: (2026)
von: Nasir, Azkaa, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Online Clustering of Dueling Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025) -
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
von: Verma, Arun, et al.
Veröffentlicht: (2024) -
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks
von: Hong, Zhi, et al.
Veröffentlicht: (2026) -
FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits
von: Lu, Pingchen, et al.
Veröffentlicht: (2025) -
Federated Linear Dueling Bandits
von: Huang, Xuhan, et al.
Veröffentlicht: (2025)