Saved in:
| Main Authors: | Biré, Emilien, Santos, María, Yuan, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.22701 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking
by: Choi, Andrew, et al.
Published: (2026)
by: Choi, Andrew, et al.
Published: (2026)
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
by: Lee, Junkyu, et al.
Published: (2025)
by: Lee, Junkyu, et al.
Published: (2025)
Process Reward Model with Q-Value Rankings
by: Li, Wendi, et al.
Published: (2024)
by: Li, Wendi, et al.
Published: (2024)
GHQ: Grouped Hybrid Q Learning for Heterogeneous Cooperative Multi-agent Reinforcement Learning
by: Yu, Xiaoyang, et al.
Published: (2023)
by: Yu, Xiaoyang, et al.
Published: (2023)
ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
by: Zhao, Kai, et al.
Published: (2023)
by: Zhao, Kai, et al.
Published: (2023)
QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
by: Lin, Zongyu, et al.
Published: (2025)
by: Lin, Zongyu, et al.
Published: (2025)
Multi-agent Reinforcement Learning with Deep Networks for Diverse Q-Vectors
by: Luo, Zhenglong, et al.
Published: (2024)
by: Luo, Zhenglong, et al.
Published: (2024)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
by: An, Selim, et al.
Published: (2026)
by: An, Selim, et al.
Published: (2026)
Stochastic Q-learning for Large Discrete Action Spaces
by: Fourati, Fares, et al.
Published: (2024)
by: Fourati, Fares, et al.
Published: (2024)
Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
by: Tian, Arther, et al.
Published: (2025)
by: Tian, Arther, et al.
Published: (2025)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
by: Ghosh, Ipsita, et al.
Published: (2025)
by: Ghosh, Ipsita, et al.
Published: (2025)
Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning
by: Wang, Chaojie, et al.
Published: (2024)
by: Wang, Chaojie, et al.
Published: (2024)
Adaptive Action Chunking via Multi-Chunk Q Value Estimation
by: Shin, Yongjae, et al.
Published: (2026)
by: Shin, Yongjae, et al.
Published: (2026)
Inference of Deterministic Finite Automata via Q-Learning
by: Hosseinkhani, Elaheh, et al.
Published: (2025)
by: Hosseinkhani, Elaheh, et al.
Published: (2025)
PersonalQ: Select, Quantize, and Serve Personalized Diffusion Models for Efficient Inference
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
by: Wu, Frank, et al.
Published: (2025)
by: Wu, Frank, et al.
Published: (2025)
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning
by: Dodeja, Lakshita, et al.
Published: (2026)
by: Dodeja, Lakshita, et al.
Published: (2026)
$β$-DQN: Improving Deep Q-Learning By Evolving the Behavior
by: Zhang, Hongming, et al.
Published: (2025)
by: Zhang, Hongming, et al.
Published: (2025)
Improving Global Parameter-sharing in Physically Heterogeneous Multi-agent Reinforcement Learning with Unified Action Space
by: Yu, Xiaoyang, et al.
Published: (2024)
by: Yu, Xiaoyang, et al.
Published: (2024)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
Causal Deep Q Network
by: Khelifi, Elouanes, et al.
Published: (2025)
by: Khelifi, Elouanes, et al.
Published: (2025)
MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs
by: Liao, Junwei, et al.
Published: (2026)
by: Liao, Junwei, et al.
Published: (2026)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Smart Sampling: Self-Attention and Bootstrapping for Improved Ensembled Q-Learning
by: Khan, Muhammad Junaid, et al.
Published: (2024)
by: Khan, Muhammad Junaid, et al.
Published: (2024)
Drift Q-Learning
by: Houssaini, Anas, et al.
Published: (2026)
by: Houssaini, Anas, et al.
Published: (2026)
Frictional Q-Learning
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Flow Q-Learning
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Decoupled Q-Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration
by: Zhang, Ziqi, et al.
Published: (2023)
by: Zhang, Ziqi, et al.
Published: (2023)
Chunk-Guided Q-Learning
by: Song, Gwanwoo, et al.
Published: (2026)
by: Song, Gwanwoo, et al.
Published: (2026)
Periodic Regularized Q-Learning
by: Yang, Hyukjun, et al.
Published: (2026)
by: Yang, Hyukjun, et al.
Published: (2026)
Deep Double Q-learning
by: Nagarajan, Prabhat, et al.
Published: (2025)
by: Nagarajan, Prabhat, et al.
Published: (2025)
SQT -- std $Q$-target
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Scalable In-Context Q-Learning
by: Liu, Jinmei, et al.
Published: (2025)
by: Liu, Jinmei, et al.
Published: (2025)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
by: Allard, Marc-Antoine, et al.
Published: (2024)
by: Allard, Marc-Antoine, et al.
Published: (2024)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
by: Singh, Aditya Kumar, et al.
Published: (2026)
by: Singh, Aditya Kumar, et al.
Published: (2026)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Techniques to Improve Q&A Accuracy with Transformer-based models on Large Complex Documents
by: Liao, Chejui, et al.
Published: (2020)
by: Liao, Chejui, et al.
Published: (2020)
Regularized Q-Learning with Linear Function Approximation
by: Xi, Jiachen, et al.
Published: (2024)
by: Xi, Jiachen, et al.
Published: (2024)
FM3Q: Factorized Multi-Agent MiniMax Q-Learning for Two-Team Zero-Sum Markov Game
by: Hu, Guangzheng, et al.
Published: (2024)
by: Hu, Guangzheng, et al.
Published: (2024)
Similar Items
-
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking
by: Choi, Andrew, et al.
Published: (2026) -
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
by: Lee, Junkyu, et al.
Published: (2025) -
Process Reward Model with Q-Value Rankings
by: Li, Wendi, et al.
Published: (2024) -
GHQ: Grouped Hybrid Q Learning for Heterogeneous Cooperative Multi-agent Reinforcement Learning
by: Yu, Xiaoyang, et al.
Published: (2023) -
ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
by: Zhao, Kai, et al.
Published: (2023)