In-Context Compositional Q-Learning for Offline Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Qiushui, Huang, Yuhao, Jiang, Yushu, Song, Lei, Wang, Jinyu, Zheng, Wenliang, Bian, Jiang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911523854614528
author Xu, Qiushui
Huang, Yuhao
Jiang, Yushu
Song, Lei
Wang, Jinyu
Zheng, Wenliang
Bian, Jiang
author_facet Xu, Qiushui
Huang, Yuhao
Jiang, Yushu
Song, Lei
Wang, Jinyu
Zheng, Wenliang
Bian, Jiang
contents Accurate estimation of the Q-function is a central challenge in offline reinforcement learning. However, existing approaches often rely on a shared global Q-function, which is inadequate for capturing the compositional structure of tasks that consist of diverse subtasks. We propose In-context Compositional Q-Learning (ICQL), an offline RL framework that formulates Q-learning as a contextual inference problem and uses linear Transformers to adaptively infer local Q-functions from retrieved transitions without explicit subtask labels. Theoretically, we show that, under two assumptions -- linear approximability of the local Q-function and accurate inference of weights from retrieved context -- ICQL achieves a bounded approximation error for the Q-function and enables near-optimal policy extraction. Empirically, ICQL substantially improves performance in offline settings, achieving gains of up to 16.4% on kitchen tasks and up to 8.8% and 6.3% on MuJoCo and Adroit tasks, respectively. These results highlight the underexplored potential of in-context learning for robust and compositional value estimation and establish ICQL as a principled and effective framework for offline RL.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24067
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle In-Context Compositional Q-Learning for Offline Reinforcement Learning
Xu, Qiushui
Huang, Yuhao
Jiang, Yushu
Song, Lei
Wang, Jinyu
Zheng, Wenliang
Bian, Jiang
Machine Learning
Artificial Intelligence
Accurate estimation of the Q-function is a central challenge in offline reinforcement learning. However, existing approaches often rely on a shared global Q-function, which is inadequate for capturing the compositional structure of tasks that consist of diverse subtasks. We propose In-context Compositional Q-Learning (ICQL), an offline RL framework that formulates Q-learning as a contextual inference problem and uses linear Transformers to adaptively infer local Q-functions from retrieved transitions without explicit subtask labels. Theoretically, we show that, under two assumptions -- linear approximability of the local Q-function and accurate inference of weights from retrieved context -- ICQL achieves a bounded approximation error for the Q-function and enables near-optimal policy extraction. Empirically, ICQL substantially improves performance in offline settings, achieving gains of up to 16.4% on kitchen tasks and up to 8.8% and 6.3% on MuJoCo and Adroit tasks, respectively. These results highlight the underexplored potential of in-context learning for robust and compositional value estimation and establish ICQL as a principled and effective framework for offline RL.
title In-Context Compositional Q-Learning for Offline Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.24067