UCB-driven Utility Function Search for Multi-objective Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yucheng, Lynch, David, Agapitos, Alexandros |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continual Model-based Reinforcement Learning for Data Efficient Wireless Network Optimisation
by: Hasan, Cengis, et al.
Published: (2024)
by: Hasan, Cengis, et al.
Published: (2024)
Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning
by: Vamplew, Peter, et al.
Published: (2024)
by: Vamplew, Peter, et al.
Published: (2024)
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2022)
by: Qiu, Shuang, et al.
Published: (2022)
Minimizing UCB: a Better Local Search Strategy in Local Bayesian Optimization
by: Fan, Zheyi, et al.
Published: (2024)
by: Fan, Zheyi, et al.
Published: (2024)
Replicable Bandits with UCB based Exploration
by: Deb, Rohan, et al.
Published: (2026)
by: Deb, Rohan, et al.
Published: (2026)
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
by: Terekhov, Mikhail, et al.
Published: (2024)
by: Terekhov, Mikhail, et al.
Published: (2024)
Extended UCB Policies for Multi-armed Bandit Problems
by: Liu, Keqin, et al.
Published: (2011)
by: Liu, Keqin, et al.
Published: (2011)
UCB-type Algorithm for Budget-Constrained Expert Learning
by: Latypov, Ilgam, et al.
Published: (2025)
by: Latypov, Ilgam, et al.
Published: (2025)
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
by: Huang, Ruiquan, et al.
Published: (2023)
by: Huang, Ruiquan, et al.
Published: (2023)
Cooperative Multi-Agent Graph Bandits: UCB Algorithm and Regret Analysis
by: Paschalidis, Phevos, et al.
Published: (2024)
by: Paschalidis, Phevos, et al.
Published: (2024)
A characterization of sample adaptivity in UCB data
by: Chen, Yilun, et al.
Published: (2025)
by: Chen, Yilun, et al.
Published: (2025)
On the Suboptimality of GP-UCB under Polynomial Effective Optimism
by: Wang, Wenjia, et al.
Published: (2023)
by: Wang, Wenjia, et al.
Published: (2023)
Truncated LinUCB for Stochastic Linear Bandits
by: Song, Yanglei, et al.
Published: (2022)
by: Song, Yanglei, et al.
Published: (2022)
UCB Exploration for Fixed-Budget Bayesian Best Arm Identification
by: Zhu, Rong J. B., et al.
Published: (2024)
by: Zhu, Rong J. B., et al.
Published: (2024)
Tractable Instances of Bilinear Maximization: Implementing LinUCB on Ellipsoids
by: Zhang, Raymond, et al.
Published: (2025)
by: Zhang, Raymond, et al.
Published: (2025)
UCB for Large-Scale Pure Exploration: Beyond Sub-Gaussianity
by: Li, Zaile, et al.
Published: (2025)
by: Li, Zaile, et al.
Published: (2025)
Clus-UCB: A Near-Optimal Algorithm for Clustered Bandits
by: Gore, Aakash, et al.
Published: (2025)
by: Gore, Aakash, et al.
Published: (2025)
Deep Multi-Objective Reinforcement Learning for Utility-Based Infrastructural Maintenance Optimization
by: van Remmerden, Jesse, et al.
Published: (2024)
by: van Remmerden, Jesse, et al.
Published: (2024)
Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs
by: Thoppe, Gugan, et al.
Published: (2026)
by: Thoppe, Gugan, et al.
Published: (2026)
Variance-Aware Linear UCB with Deep Representation for Neural Contextual Bandits
by: Bui, Ha Manh, et al.
Published: (2024)
by: Bui, Ha Manh, et al.
Published: (2024)
Revisiting Social Welfare in Bandits: UCB is (Nearly) All You Need
by: Sarkar, Dhruv, et al.
Published: (2025)
by: Sarkar, Dhruv, et al.
Published: (2025)
DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative Models
by: Jafari, Donya, et al.
Published: (2026)
by: Jafari, Donya, et al.
Published: (2026)
Precise Asymptotics and Refined Regret of Variance-Aware UCB
by: Fan, Yingying, et al.
Published: (2024)
by: Fan, Yingying, et al.
Published: (2024)
PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs
by: Hu, Xiaoyan, et al.
Published: (2024)
by: Hu, Xiaoyan, et al.
Published: (2024)
Multi-Target Radar Search and Track Using Sequence-Capable Deep Reinforcement Learning
by: Ewers, Jan-Hendrik, et al.
Published: (2025)
by: Ewers, Jan-Hendrik, et al.
Published: (2025)
Modeling Multi-Objective Tradeoffs with Monotonic Utility Functions
by: Chen, Edward, et al.
Published: (2024)
by: Chen, Edward, et al.
Published: (2024)
A Spatially Informed Gaussian Process UCB Method for Decentralized Coverage Control
by: Guidone, Gennaro, et al.
Published: (2025)
by: Guidone, Gennaro, et al.
Published: (2025)
Issues with Value-Based Multi-objective Reinforcement Learning: Value Function Interference and Overestimation Sensitivity
by: Vamplew, Peter, et al.
Published: (2024)
by: Vamplew, Peter, et al.
Published: (2024)
Statistical Inference under Adaptive Sampling with LinUCB
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
Polynomial Regret Concentration of UCB for Non-Deterministic State Transitions
by: Cömer, Can, et al.
Published: (2025)
by: Cömer, Can, et al.
Published: (2025)
Reward-Based Online LLM Routing via NeuralUCB
by: Tsai, Ming-Hua, et al.
Published: (2026)
by: Tsai, Ming-Hua, et al.
Published: (2026)
On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
by: Dong, Zixuan, et al.
Published: (2022)
by: Dong, Zixuan, et al.
Published: (2022)
Deep Reinforcement Learning for Ranking Utility Tuning in the Ad Recommender System at Pinterest
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
Soft Forward-Backward Representations for Zero-shot Reinforcement Learning with General Utilities
by: Bagatella, Marco, et al.
Published: (2026)
by: Bagatella, Marco, et al.
Published: (2026)
Federated Learning for Data Market: Shapley-UCB for Seller Selection and Incentives
by: Chen, Kongyang, et al.
Published: (2024)
by: Chen, Kongyang, et al.
Published: (2024)
Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter Efficient
by: Wang, Wenlong, et al.
Published: (2024)
by: Wang, Wenlong, et al.
Published: (2024)
Trajectory Modeling via Random Utility Inverse Reinforcement Learning
by: Pitombeira-Neto, Anselmo R., et al.
Published: (2021)
by: Pitombeira-Neto, Anselmo R., et al.
Published: (2021)
Search Inspired Exploration in Reinforcement Learning
by: Sotirchos, Georgios, et al.
Published: (2026)
by: Sotirchos, Georgios, et al.
Published: (2026)
Privacy-Preserving UCB Decision Process Verification via zk-SNARKs
by: Jiang, Xikun, et al.
Published: (2024)
by: Jiang, Xikun, et al.
Published: (2024)
Efficient Implementation of LinearUCB through Algorithmic Improvements and Vector Computing Acceleration for Embedded Learning Systems
by: Angioli, Marco, et al.
Published: (2025)
by: Angioli, Marco, et al.
Published: (2025)
Similar Items
-
Continual Model-based Reinforcement Learning for Data Efficient Wireless Network Optimisation
by: Hasan, Cengis, et al.
Published: (2024) -
Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning
by: Vamplew, Peter, et al.
Published: (2024) -
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2022) -
Minimizing UCB: a Better Local Search Strategy in Local Bayesian Optimization
by: Fan, Zheyi, et al.
Published: (2024) -
Replicable Bandits with UCB based Exploration
by: Deb, Rohan, et al.
Published: (2026)