Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Lang, Tan, Weihao, Lyu, Zhiyi, Zheng, Longtao, Xu, Haiyang, Yan, Ming, Huang, Fei, An, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024)
by: Tan, Weihao, et al.
Published: (2024)
AgentOCR: Reimagining Agent History via Optical Self-Compression
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
AgentStudio: A Toolkit for Building General Virtual Agents
by: Zheng, Longtao, et al.
Published: (2024)
by: Zheng, Longtao, et al.
Published: (2024)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
by: Jiang, Chaoya, et al.
Published: (2025)
by: Jiang, Chaoya, et al.
Published: (2025)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
by: Wang, Junyang, et al.
Published: (2025)
by: Wang, Junyang, et al.
Published: (2025)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
by: Wang, Junyang, et al.
Published: (2025)
by: Wang, Junyang, et al.
Published: (2025)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
by: Lu, Zhengxi, et al.
Published: (2025)
by: Lu, Zhengxi, et al.
Published: (2025)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
CoINS: Counterfactual Interactive Navigation via Skill-Aware VLM
by: Zhou, Kangjie, et al.
Published: (2026)
by: Zhou, Kangjie, et al.
Published: (2026)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies
by: Chen, Keyu, et al.
Published: (2026)
by: Chen, Keyu, et al.
Published: (2026)
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
by: Gao, Zhi, et al.
Published: (2024)
by: Gao, Zhi, et al.
Published: (2024)
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
by: Wu, Xiefeng, et al.
Published: (2025)
by: Wu, Xiefeng, et al.
Published: (2025)
FishBargain: An LLM-Empowered Bargaining Agent for Online Fleamarket Platform Sellers
by: Kong, Dexin, et al.
Published: (2025)
by: Kong, Dexin, et al.
Published: (2025)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
by: Wang, Kangrui, et al.
Published: (2025)
by: Wang, Kangrui, et al.
Published: (2025)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
by: Song, Feifan, et al.
Published: (2024)
by: Song, Feifan, et al.
Published: (2024)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
TRELM: Towards Robust and Efficient Pre-training for Knowledge-Enhanced Language Models
by: Yan, Junbing, et al.
Published: (2024)
by: Yan, Junbing, et al.
Published: (2024)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
by: Zhang, Liang, et al.
Published: (2024)
by: Zhang, Liang, et al.
Published: (2024)
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
by: Xue, Zhenghai, et al.
Published: (2025)
by: Xue, Zhenghai, et al.
Published: (2025)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
by: Zhang, Yimeng, et al.
Published: (2025)
by: Zhang, Yimeng, et al.
Published: (2025)
LBM: Hierarchical Large Auto-Bidding Model via Reasoning and Acting
by: Li, Yewen, et al.
Published: (2026)
by: Li, Yewen, et al.
Published: (2026)
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning
by: Wu, Qingyuan, et al.
Published: (2025)
by: Wu, Qingyuan, et al.
Published: (2025)
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
by: Zheng, Longtao, et al.
Published: (2023)
by: Zheng, Longtao, et al.
Published: (2023)
Extend Model Merging from Fine-Tuned to Pre-Trained Large Language Models via Weight Disentanglement
by: Yu, Le, et al.
Published: (2024)
by: Yu, Le, et al.
Published: (2024)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
by: Hu, Xuhao, et al.
Published: (2026)
by: Hu, Xuhao, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
SCENE: Evaluating Explainable AI Techniques Using Soft Counterfactuals
by: Zheng, Haoran, et al.
Published: (2024)
by: Zheng, Haoran, et al.
Published: (2024)
Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
by: Wang, Wentao, et al.
Published: (2025)
by: Wang, Wentao, et al.
Published: (2025)
Sample-Efficient Counterfactual Tuning for Compressor Pressure Control
by: Guerrero, Margarita A., et al.
Published: (2025)
by: Guerrero, Margarita A., et al.
Published: (2025)
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
by: Lei, Qinqian, et al.
Published: (2024)
by: Lei, Qinqian, et al.
Published: (2024)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025)
by: Zhang, Chubin, et al.
Published: (2025)
Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs
by: Lyu, Zhiyi, et al.
Published: (2025)
by: Lyu, Zhiyi, et al.
Published: (2025)
Similar Items
-
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026) -
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024) -
AgentOCR: Reimagining Agent History via Optical Self-Compression
by: Feng, Lang, et al.
Published: (2026) -
AgentStudio: A Toolkit for Building General Virtual Agents
by: Zheng, Longtao, et al.
Published: (2024) -
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
by: Jiang, Chaoya, et al.
Published: (2025)