Q-Pensieve: Boosting Sample Efficiency of Multi-Objective RL Through Memory Sharing of Q-Snapshots
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hung, Wei, Huang, Bo-Kai, Hsieh, Ping-Chun, Liu, Xi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025)
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025)
Enhancing Offline Model-Based RL via Active Model Selection: A Bayesian Optimization Perspective
von: Yang, Yu-Wei, et al.
Veröffentlicht: (2025)
von: Yang, Yu-Wei, et al.
Veröffentlicht: (2025)
A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning
von: Chen, Ying-Tu, et al.
Veröffentlicht: (2026)
von: Chen, Ying-Tu, et al.
Veröffentlicht: (2026)
Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025)
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention
von: Chen, Yuxin, et al.
Veröffentlicht: (2024)
von: Chen, Yuxin, et al.
Veröffentlicht: (2024)
Surrogate Ensemble in Expensive Multi-Objective Optimization via Deep Q-Learning
von: Wu, Yuxin, et al.
Veröffentlicht: (2026)
von: Wu, Yuxin, et al.
Veröffentlicht: (2026)
Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
von: Guo, Jian-Ting, et al.
Veröffentlicht: (2025)
von: Guo, Jian-Ting, et al.
Veröffentlicht: (2025)
Q-learning with Posterior Sampling
von: Agrawal, Priyank, et al.
Veröffentlicht: (2025)
von: Agrawal, Priyank, et al.
Veröffentlicht: (2025)
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
von: Bhatt, Aditya, et al.
Veröffentlicht: (2019)
von: Bhatt, Aditya, et al.
Veröffentlicht: (2019)
QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
von: Lin, Zongyu, et al.
Veröffentlicht: (2025)
von: Lin, Zongyu, et al.
Veröffentlicht: (2025)
Boosting Soft Q-Learning by Bounding
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
Stateful Large Language Model Serving with Pensieve
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
Yes, Q-learning Helps Offline In-Context RL
von: Tarasov, Denis, et al.
Veröffentlicht: (2025)
von: Tarasov, Denis, et al.
Veröffentlicht: (2025)
Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPs
von: Hung, Wei, et al.
Veröffentlicht: (2025)
von: Hung, Wei, et al.
Veröffentlicht: (2025)
QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing
von: Zhang, Grace, et al.
Veröffentlicht: (2023)
von: Zhang, Grace, et al.
Veröffentlicht: (2023)
Asymptotic Analysis of Sample-averaged Q-learning
von: Panda, Saunak Kumar, et al.
Veröffentlicht: (2024)
von: Panda, Saunak Kumar, et al.
Veröffentlicht: (2024)
Multi-Actor Multi-Critic Deep Deterministic Reinforcement Learning with a Novel Q-Ensemble Method
von: Wu, Andy, et al.
Veröffentlicht: (2025)
von: Wu, Andy, et al.
Veröffentlicht: (2025)
eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization
von: Su, Pei-Chun
Veröffentlicht: (2026)
von: Su, Pei-Chun
Veröffentlicht: (2026)
Boosting Sample Efficiency and Generalization in Multi-agent Reinforcement Learning via Equivariance
von: McClellan, Joshua, et al.
Veröffentlicht: (2024)
von: McClellan, Joshua, et al.
Veröffentlicht: (2024)
Behavior-Adaptive Q-Learning: A Unifying Framework for Offline-to-Online RL
von: Zu, Lipeng, et al.
Veröffentlicht: (2025)
von: Zu, Lipeng, et al.
Veröffentlicht: (2025)
Diminishing Exploration: A Minimalist Approach to Piecewise Stationary Multi-Armed Bandits
von: Li, Kuan-Ta, et al.
Veröffentlicht: (2024)
von: Li, Kuan-Ta, et al.
Veröffentlicht: (2024)
Uniformly Safe RL with Objective Suppression for Multi-Constraint Safety-Critical Applications
von: Zhou, Zihan, et al.
Veröffentlicht: (2024)
von: Zhou, Zihan, et al.
Veröffentlicht: (2024)
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
von: Wang, Shengbo
Veröffentlicht: (2026)
von: Wang, Shengbo
Veröffentlicht: (2026)
Missing Data Multiple Imputation for Tabular Q-Learning in Online RL
von: Chasalow, Kyla, et al.
Veröffentlicht: (2025)
von: Chasalow, Kyla, et al.
Veröffentlicht: (2025)
QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL
von: Lei, Xing, et al.
Veröffentlicht: (2026)
von: Lei, Xing, et al.
Veröffentlicht: (2026)
Enhancing Sample Efficiency in Multi-Agent RL with Uncertainty Quantification and Selective Exploration
von: Danino, Tom, et al.
Veröffentlicht: (2025)
von: Danino, Tom, et al.
Veröffentlicht: (2025)
Pruning the Way to Reliable Policies: A Multi-Objective Deep Q-Learning Approach to Critical Care
von: Shirali, Ali, et al.
Veröffentlicht: (2023)
von: Shirali, Ali, et al.
Veröffentlicht: (2023)
Snapshot Reinforcement Learning: Leveraging Prior Trajectories for Efficiency
von: Zhao, Yanxiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yanxiao, et al.
Veröffentlicht: (2024)
Cross-Domain Policy Optimization via Bellman Consistency and Hybrid Critics
von: Chen, Ming-Hong, et al.
Veröffentlicht: (2026)
von: Chen, Ming-Hong, et al.
Veröffentlicht: (2026)
Scalable In-Context Q-Learning
von: Liu, Jinmei, et al.
Veröffentlicht: (2025)
von: Liu, Jinmei, et al.
Veröffentlicht: (2025)
Q-WSL: Optimizing Goal-Conditioned RL with Weighted Supervised Learning via Dynamic Programming
von: Lei, Xing, et al.
Veröffentlicht: (2024)
von: Lei, Xing, et al.
Veröffentlicht: (2024)
Maximizing Data Efficiency for Cross-Lingual TTS Adaptation by Self-Supervised Representation Mixing and Embedding Initialization
von: Huang, Wei-Ping, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Ping, et al.
Veröffentlicht: (2024)
Data-Driven Knowledge Transfer in Batch $Q^*$ Learning
von: Chen, Elynn, et al.
Veröffentlicht: (2024)
von: Chen, Elynn, et al.
Veröffentlicht: (2024)
SDM-Q: Cost-Aware Staged Decision-Making for Multi-Omics Classification with Deep Q-Learning
von: Mu, Nan, et al.
Veröffentlicht: (2026)
von: Mu, Nan, et al.
Veröffentlicht: (2026)
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2026)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2026)
HyperQ-Opt: Q-learning for Hyperparameter Optimization
von: Hasan, Md. Tarek
Veröffentlicht: (2024)
von: Hasan, Md. Tarek
Veröffentlicht: (2024)
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
von: An, Selim, et al.
Veröffentlicht: (2026)
von: An, Selim, et al.
Veröffentlicht: (2026)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025) -
Enhancing Offline Model-Based RL via Active Model Selection: A Bayesian Optimization Perspective
von: Yang, Yu-Wei, et al.
Veröffentlicht: (2025) -
A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning
von: Chen, Ying-Tu, et al.
Veröffentlicht: (2026) -
Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025) -
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)