Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Lang, Tan, Weihao, Lyu, Zhiyi, Zheng, Longtao, Xu, Haiyang, Yan, Ming, Huang, Fei, An, Bo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915319515185152
author Feng, Lang
Tan, Weihao
Lyu, Zhiyi
Zheng, Longtao
Xu, Haiyang
Yan, Ming
Huang, Fei
An, Bo
author_facet Feng, Lang
Tan, Weihao
Lyu, Zhiyi
Zheng, Longtao
Xu, Haiyang
Yan, Ming
Huang, Fei
An, Bo
contents Online fine-tuning vision-language model (VLM) agents with reinforcement learning (RL) has shown promise for equipping agents with multi-step, goal-oriented capabilities in dynamic environments. However, their open-ended textual action space and non-end-to-end nature of action generation present significant challenges to effective online exploration in RL, e.g., explosion of the exploration space. We propose a novel online fine-tuning method, Counterfactual Soft Reinforcement Learning (CoSo), better suited to the textual output space of VLM agents. Compared to prior methods that assign uniform uncertainty to all tokens, CoSo leverages counterfactual reasoning to dynamically assess the causal influence of individual tokens on post-processed actions. By prioritizing the exploration of action-critical tokens while reducing the impact of semantically redundant or low-impact tokens, CoSo enables a more targeted and efficient online rollout process. We provide theoretical analysis proving CoSo's convergence and policy improvement guarantees, and extensive empirical evaluations supporting CoSo's effectiveness. Our results across a diverse set of agent tasks, including Android device control, card gaming, and embodied AI, highlight its remarkable ability to enhance exploration efficiency and deliver consistent performance gains. The code is available at https://github.com/langfengQ/CoSo.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
Feng, Lang
Tan, Weihao
Lyu, Zhiyi
Zheng, Longtao
Xu, Haiyang
Yan, Ming
Huang, Fei
An, Bo
Machine Learning
Artificial Intelligence
Online fine-tuning vision-language model (VLM) agents with reinforcement learning (RL) has shown promise for equipping agents with multi-step, goal-oriented capabilities in dynamic environments. However, their open-ended textual action space and non-end-to-end nature of action generation present significant challenges to effective online exploration in RL, e.g., explosion of the exploration space. We propose a novel online fine-tuning method, Counterfactual Soft Reinforcement Learning (CoSo), better suited to the textual output space of VLM agents. Compared to prior methods that assign uniform uncertainty to all tokens, CoSo leverages counterfactual reasoning to dynamically assess the causal influence of individual tokens on post-processed actions. By prioritizing the exploration of action-critical tokens while reducing the impact of semantically redundant or low-impact tokens, CoSo enables a more targeted and efficient online rollout process. We provide theoretical analysis proving CoSo's convergence and policy improvement guarantees, and extensive empirical evaluations supporting CoSo's effectiveness. Our results across a diverse set of agent tasks, including Android device control, card gaming, and embodied AI, highlight its remarkable ability to enhance exploration efficiency and deliver consistent performance gains. The code is available at https://github.com/langfengQ/CoSo.
title Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.03792