Continuous K-Max Bandits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yu, Wang, Siwei, Huang, Longbo, Chen, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
von: Chen, Yu, et al.
Veröffentlicht: (2024)
von: Chen, Yu, et al.
Veröffentlicht: (2024)
Combinatorial Rising Bandits
von: Song, Seockbean, et al.
Veröffentlicht: (2024)
von: Song, Seockbean, et al.
Veröffentlicht: (2024)
Adversarial Network Optimization under Bandit Feedback: Maximizing Utility in Non-Stationary Multi-Hop Networks
von: Dai, Yan, et al.
Veröffentlicht: (2024)
von: Dai, Yan, et al.
Veröffentlicht: (2024)
Provably Efficient Partially Observable Risk-Sensitive Reinforcement Learning with Hindsight Observation
von: Zhang, Tonghe, et al.
Veröffentlicht: (2024)
von: Zhang, Tonghe, et al.
Veröffentlicht: (2024)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs
von: Chen, Yu, et al.
Veröffentlicht: (2024)
von: Chen, Yu, et al.
Veröffentlicht: (2024)
Rising Multi-Armed Bandits with Known Horizons
von: Song, Seockbean, et al.
Veröffentlicht: (2026)
von: Song, Seockbean, et al.
Veröffentlicht: (2026)
Materials Discovery using Max K-Armed Bandit
von: Kikkawa, Nobuaki, et al.
Veröffentlicht: (2022)
von: Kikkawa, Nobuaki, et al.
Veröffentlicht: (2022)
Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants
von: Liu, Yuhao, et al.
Veröffentlicht: (2025)
von: Liu, Yuhao, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Discrete-Time Stochastic Interpolants
von: Liu, Yuhao, et al.
Veröffentlicht: (2025)
von: Liu, Yuhao, et al.
Veröffentlicht: (2025)
Offline Learning for Combinatorial Multi-armed Bandits
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes
von: Chen, Yu, et al.
Veröffentlicht: (2026)
von: Chen, Yu, et al.
Veröffentlicht: (2026)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
von: Chen, Ruishuo, et al.
Veröffentlicht: (2026)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2026)
Decentralized Asynchronous Multi-player Bandits
von: Fan, Jingqi, et al.
Veröffentlicht: (2025)
von: Fan, Jingqi, et al.
Veröffentlicht: (2025)
Bandit Max-Min Fair Allocation
von: Harada, Tsubasa, et al.
Veröffentlicht: (2025)
von: Harada, Tsubasa, et al.
Veröffentlicht: (2025)
Learning with Limited Shared Information in Multi-agent Multi-armed Bandit
von: Shao, Junning, et al.
Veröffentlicht: (2025)
von: Shao, Junning, et al.
Veröffentlicht: (2025)
Efficient and Optimal Policy Gradient Algorithm for Corrupted Multi-armed Bandits
von: Liu, Jiayuan, et al.
Veröffentlicht: (2025)
von: Liu, Jiayuan, et al.
Veröffentlicht: (2025)
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
von: Liu, Xutong, et al.
Veröffentlicht: (2023)
von: Liu, Xutong, et al.
Veröffentlicht: (2023)
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
Mixed Sparsity Training: Achieving 4$\times$ FLOP Reduction for Transformer Pretraining
von: Hu, Pihe, et al.
Veröffentlicht: (2024)
von: Hu, Pihe, et al.
Veröffentlicht: (2024)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
von: Liu, Xutong, et al.
Veröffentlicht: (2022)
von: Liu, Xutong, et al.
Veröffentlicht: (2022)
Achieving Optimal Static and Dynamic Regret Simultaneously in Bandits with Deterministic Losses
von: Qian, Jian, et al.
Veröffentlicht: (2026)
von: Qian, Jian, et al.
Veröffentlicht: (2026)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
Collaborative Min-Max Regret in Grouped Multi-Armed Bandits
von: Blanchard, Moïse, et al.
Veröffentlicht: (2025)
von: Blanchard, Moïse, et al.
Veröffentlicht: (2025)
Real-Time Parallel Counterfactual Regret Minimization
von: Li, Boning, et al.
Veröffentlicht: (2026)
von: Li, Boning, et al.
Veröffentlicht: (2026)
Layer-Aware Influence for Online Data Valuation Estimation
von: Yang, Ziao, et al.
Veröffentlicht: (2025)
von: Yang, Ziao, et al.
Veröffentlicht: (2025)
How Does Variance Shape the Regret in Contextual Bandits?
von: Jia, Zeyu, et al.
Veröffentlicht: (2024)
von: Jia, Zeyu, et al.
Veröffentlicht: (2024)
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
Reparameterization Flow Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
Multi-Play Combinatorial Semi-Bandit Problem
von: Nakamura, Shintaro, et al.
Veröffentlicht: (2025)
von: Nakamura, Shintaro, et al.
Veröffentlicht: (2025)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
von: van Erven, Tim, et al.
Veröffentlicht: (2025)
von: van Erven, Tim, et al.
Veröffentlicht: (2025)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
RL-CFR: Improving Action Abstraction for Imperfect Information Extensive-Form Games with Reinforcement Learning
von: Li, Boning, et al.
Veröffentlicht: (2024)
von: Li, Boning, et al.
Veröffentlicht: (2024)
Continuous Semantic Caching for Low-Cost LLM Serving
von: Atalar, Baran, et al.
Veröffentlicht: (2026)
von: Atalar, Baran, et al.
Veröffentlicht: (2026)
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
von: Wang, Xun, et al.
Veröffentlicht: (2025)
von: Wang, Xun, et al.
Veröffentlicht: (2025)
Diversity-Preserving K-Armed Bandits, Revisited
von: Hadiji, Hédi, et al.
Veröffentlicht: (2020)
von: Hadiji, Hédi, et al.
Veröffentlicht: (2020)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
von: Hu, Rui, et al.
Veröffentlicht: (2024)
von: Hu, Rui, et al.
Veröffentlicht: (2024)
A Quadratic Synchronization Rule for Distributed Deep Learning
von: Gu, Xinran, et al.
Veröffentlicht: (2023)
von: Gu, Xinran, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
von: Chen, Yu, et al.
Veröffentlicht: (2024) -
Combinatorial Rising Bandits
von: Song, Seockbean, et al.
Veröffentlicht: (2024) -
Adversarial Network Optimization under Bandit Feedback: Maximizing Utility in Non-Stationary Multi-Hop Networks
von: Dai, Yan, et al.
Veröffentlicht: (2024) -
Provably Efficient Partially Observable Risk-Sensitive Reinforcement Learning with Hindsight Observation
von: Zhang, Tonghe, et al.
Veröffentlicht: (2024) -
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
von: Hu, Rui, et al.
Veröffentlicht: (2025)