rQdia: Regularizing Q-Value Distributions With Image Augmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Lerman, Sam, Bi, Jing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Periodic Regularized Q-Learning
por: Yang, Hyukjun, et al.
Publicado: (2026)
por: Yang, Hyukjun, et al.
Publicado: (2026)
Adaptive Action Chunking via Multi-Chunk Q Value Estimation
por: Shin, Yongjae, et al.
Publicado: (2026)
por: Shin, Yongjae, et al.
Publicado: (2026)
Residual Q-Learning: Offline and Online Policy Customization without Value
por: Li, Chenran, et al.
Publicado: (2023)
por: Li, Chenran, et al.
Publicado: (2023)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
por: Liu, Kezhao, et al.
Publicado: (2025)
por: Liu, Kezhao, et al.
Publicado: (2025)
Quantile Geometry Regularization for Distributional Reinforcement Learning
por: Zhang, Zhaofan, et al.
Publicado: (2026)
por: Zhang, Zhaofan, et al.
Publicado: (2026)
Towards Adapting Reinforcement Learning Agents to New Tasks: Insights from Q-Values
por: Ramaswamy, Ashwin, et al.
Publicado: (2024)
por: Ramaswamy, Ashwin, et al.
Publicado: (2024)
Value-Distributional Model-Based Reinforcement Learning
por: Luis, Carlos E., et al.
Publicado: (2023)
por: Luis, Carlos E., et al.
Publicado: (2023)
Addressing Label Shift in Distributed Learning via Entropy Regularization
por: Wu, Zhiyuan, et al.
Publicado: (2025)
por: Wu, Zhiyuan, et al.
Publicado: (2025)
Continuous Control Reinforcement Learning: Distributed Distributional DrQ Algorithms
por: Zhou, Zehao
Publicado: (2024)
por: Zhou, Zehao
Publicado: (2024)
Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples
por: Meng, Li, et al.
Publicado: (2021)
por: Meng, Li, et al.
Publicado: (2021)
Policy Regularized Distributionally Robust Markov Decision Processes with Linear Function Approximation
por: Gu, Jingwen, et al.
Publicado: (2025)
por: Gu, Jingwen, et al.
Publicado: (2025)
Per-Domain Generalizing Policies: On Learning Efficient and Robust Q-Value Functions (Extended Version with Technical Appendix)
por: Müller, Nicola J., et al.
Publicado: (2026)
por: Müller, Nicola J., et al.
Publicado: (2026)
Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies
por: Zhang, Mingming, et al.
Publicado: (2026)
por: Zhang, Mingming, et al.
Publicado: (2026)
Data-Driven Estimation of Heterogeneous Treatment Effects
por: Tran, Christopher, et al.
Publicado: (2023)
por: Tran, Christopher, et al.
Publicado: (2023)
Graph Data Augmentation with Contrastive Learning on Covariate Distribution Shift
por: Zeng, Fanlong, et al.
Publicado: (2025)
por: Zeng, Fanlong, et al.
Publicado: (2025)
ACCon: Angle-Compensated Contrastive Regularizer for Deep Regression
por: Zhao, Botao, et al.
Publicado: (2025)
por: Zhao, Botao, et al.
Publicado: (2025)
Posts of Peril: Detecting Information About Hazards in Text
por: Burghardt, Keith, et al.
Publicado: (2024)
por: Burghardt, Keith, et al.
Publicado: (2024)
From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
por: Dang, Sizhe, et al.
Publicado: (2026)
por: Dang, Sizhe, et al.
Publicado: (2026)
Distributed Multi-Agent Reinforcement Learning Based on Graph-Induced Local Value Functions
por: Jing, Gangshan, et al.
Publicado: (2022)
por: Jing, Gangshan, et al.
Publicado: (2022)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
por: Pang, Jing-Cheng, et al.
Publicado: (2021)
por: Pang, Jing-Cheng, et al.
Publicado: (2021)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
por: Ghosh, Ipsita, et al.
Publicado: (2025)
por: Ghosh, Ipsita, et al.
Publicado: (2025)
Score-based Conditional Out-of-Distribution Augmentation for Graph Covariate Shift
por: Wang, Bohan, et al.
Publicado: (2024)
por: Wang, Bohan, et al.
Publicado: (2024)
A Simple Data Augmentation for Feature Distribution Skewed Federated Learning
por: Yan, Yunlu, et al.
Publicado: (2023)
por: Yan, Yunlu, et al.
Publicado: (2023)
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
por: Wang, Qingjun, et al.
Publicado: (2026)
por: Wang, Qingjun, et al.
Publicado: (2026)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
por: Zhu, Dingwei, et al.
Publicado: (2025)
por: Zhu, Dingwei, et al.
Publicado: (2025)
Value Augmented Sampling for Language Model Alignment and Personalization
por: Han, Seungwook, et al.
Publicado: (2024)
por: Han, Seungwook, et al.
Publicado: (2024)
Kernel-Based Distributed Q-Learning: A Scalable Reinforcement Learning Approach for Dynamic Treatment Regimes
por: Wang, Di, et al.
Publicado: (2023)
por: Wang, Di, et al.
Publicado: (2023)
Learning Regularizers: Learning Optimizers that can Regularize
por: Sahoo, Suraj Kumar, et al.
Publicado: (2025)
por: Sahoo, Suraj Kumar, et al.
Publicado: (2025)
RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value Factorization
por: Shen, Siqi, et al.
Publicado: (2023)
por: Shen, Siqi, et al.
Publicado: (2023)
Efficient Post-Training Augmentation for Adaptive Inference in Heterogeneous and Distributed IoT Environments
por: Sponner, Max, et al.
Publicado: (2024)
por: Sponner, Max, et al.
Publicado: (2024)
LLM-Augmented Computational Phenotyping of Long Covid
por: Wang, Jing, et al.
Publicado: (2026)
por: Wang, Jing, et al.
Publicado: (2026)
Joint Distribution-Informed Shapley Values for Sparse Counterfactual Explanations
por: You, Lei, et al.
Publicado: (2024)
por: You, Lei, et al.
Publicado: (2024)
A Novel Spatiotemporal Coupling Graph Convolutional Network
por: Bi, Fanghui
Publicado: (2024)
por: Bi, Fanghui
Publicado: (2024)
Improve Value Estimation of Q Function and Reshape Reward with Monte Carlo Tree Search
por: Li, Jiamian
Publicado: (2024)
por: Li, Jiamian
Publicado: (2024)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
por: Guo, Hanze, et al.
Publicado: (2025)
por: Guo, Hanze, et al.
Publicado: (2025)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
por: Cho, Taehyun, et al.
Publicado: (2024)
por: Cho, Taehyun, et al.
Publicado: (2024)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
por: Lee, Jaehyeok, et al.
Publicado: (2026)
por: Lee, Jaehyeok, et al.
Publicado: (2026)
Offline Imitation Learning with Model-based Reverse Augmentation
por: Shao, Jie-Jing, et al.
Publicado: (2024)
por: Shao, Jie-Jing, et al.
Publicado: (2024)
TBBC: Predict True Bacteraemia in Blood Cultures via Deep Learning
por: Sam, Kira
Publicado: (2024)
por: Sam, Kira
Publicado: (2024)
Frictional Q-Learning
por: Kim, Hyunwoo, et al.
Publicado: (2025)
por: Kim, Hyunwoo, et al.
Publicado: (2025)
Ejemplares similares
-
Periodic Regularized Q-Learning
por: Yang, Hyukjun, et al.
Publicado: (2026) -
Adaptive Action Chunking via Multi-Chunk Q Value Estimation
por: Shin, Yongjae, et al.
Publicado: (2026) -
Residual Q-Learning: Offline and Online Policy Customization without Value
por: Li, Chenran, et al.
Publicado: (2023) -
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
por: Liu, Kezhao, et al.
Publicado: (2025) -
Quantile Geometry Regularization for Distributional Reinforcement Learning
por: Zhang, Zhaofan, et al.
Publicado: (2026)