GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiongbin, Luo, Zhihao, Lei, Shanzhe, Zhang, Lechao, Wang, Xuhong, Yang, Jie, Zheng, Zhonglong, Zheng, Yuanjie, Tan, Xin, Liu, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
by: Yuan, Yi, et al.
Published: (2026)
by: Yuan, Yi, et al.
Published: (2026)
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
by: Zhang, Lechao, et al.
Published: (2026)
by: Zhang, Lechao, et al.
Published: (2026)
CredID: Credible Multi-Bit Watermark for Large Language Models Identification
by: Jiang, Haoyu, et al.
Published: (2024)
by: Jiang, Haoyu, et al.
Published: (2024)
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
by: Zhao, Pu, et al.
Published: (2024)
by: Zhao, Pu, et al.
Published: (2024)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
AgentAlign: Misalignment-Adapted Multi-Agent Perception for Resilient Inter-Agent Sensor Correlations
by: Meng, Zonglin, et al.
Published: (2024)
by: Meng, Zonglin, et al.
Published: (2024)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Radiation‐induced fibrosis: Mechanisms and therapeutic strategies from an immune microenvironment perspective
by: Mengting Zheng, et al.
Published: (2024)
by: Mengting Zheng, et al.
Published: (2024)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO
by: Chen, Tianyang, et al.
Published: (2026)
by: Chen, Tianyang, et al.
Published: (2026)
Robust Physics-Guided Diffusion for Full-Waveform Inversion
by: Peng, Jishen, et al.
Published: (2026)
by: Peng, Jishen, et al.
Published: (2026)
Dialogue Model Optimization via Agent Game and Adaptive Tree-based GRPO
by: Peng, Kun, et al.
Published: (2026)
by: Peng, Kun, et al.
Published: (2026)
UniMark: Artificial Intelligence Generated Content Identification Toolkit
by: Li, Meilin, et al.
Published: (2025)
by: Li, Meilin, et al.
Published: (2025)
CORAL: COntextual Reasoning And Local Planning in A Hierarchical VLM Framework for Underwater Monitoring
by: Wu, Zhenqi, et al.
Published: (2026)
by: Wu, Zhenqi, et al.
Published: (2026)
CONTRIBUTE NOW, GROW TOMORROW
by: Reihan Yulizar Pratama, et al.
Published: (2026)
by: Reihan Yulizar Pratama, et al.
Published: (2026)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2026)
by: Ma, Xiaoxiao, et al.
Published: (2026)
Evaluating the fate and variability of soil organic carbon and nitrogen species under conservation practices in the Raccoon River Watershed
by: Zhonglong Zhang, et al.
Published: (2026)
by: Zhonglong Zhang, et al.
Published: (2026)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
PhysiAgent: An Embodied Agent Framework in Physical World
by: Wang, Zhihao, et al.
Published: (2025)
by: Wang, Zhihao, et al.
Published: (2025)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Odyssey: Empowering Minecraft Agents with Open-World Skills
by: Liu, Shunyu, et al.
Published: (2024)
by: Liu, Shunyu, et al.
Published: (2024)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
WDformer: A Wavelet-based Differential Transformer Model for Time Series Forecasting
by: Wang, Xiaojian, et al.
Published: (2025)
by: Wang, Xiaojian, et al.
Published: (2025)
SRSNetwork: Siamese Reconstruction-Segmentation Networks based on Dynamic-Parameter Convolution
by: Nian, Bingkun, et al.
Published: (2023)
by: Nian, Bingkun, et al.
Published: (2023)
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
by: Wu, Jiayi, et al.
Published: (2026)
by: Wu, Jiayi, et al.
Published: (2026)
Aligning VLM Assistants with Personalized Situated Cognition
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
Learning Counterfactually Decoupled Attention for Open-World Model Attribution
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
by: Wu, Xiefeng, et al.
Published: (2025)
by: Wu, Xiefeng, et al.
Published: (2025)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
by: Wang, Kangrui, et al.
Published: (2025)
by: Wang, Kangrui, et al.
Published: (2025)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
by: Yang, Zixuan, et al.
Published: (2026)
by: Yang, Zixuan, et al.
Published: (2026)
HOW DO BRIGHTEST CLUSTER GALAXIES GROW?
by: P. Oliva-Altamirano
Published: (2014)
by: P. Oliva-Altamirano
Published: (2014)
VLM-3D:End-to-End Vision-Language Models for Open-World 3D Perception
by: Chang, Fuhao, et al.
Published: (2025)
by: Chang, Fuhao, et al.
Published: (2025)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
by: Zhong, Haitian, et al.
Published: (2026)
by: Zhong, Haitian, et al.
Published: (2026)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
by: Wang, Pan, et al.
Published: (2026)
by: Wang, Pan, et al.
Published: (2026)
Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling
by: Xiao, Lechao
Published: (2024)
by: Xiao, Lechao
Published: (2024)
RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
by: Wang, Shengyuan, et al.
Published: (2025)
by: Wang, Shengyuan, et al.
Published: (2025)
Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Similar Items
-
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
by: Yuan, Yi, et al.
Published: (2026) -
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
by: Zhang, Lechao, et al.
Published: (2026) -
CredID: Credible Multi-Bit Watermark for Large Language Models Identification
by: Jiang, Haoyu, et al.
Published: (2024) -
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
by: He, Qianxi, et al.
Published: (2025) -
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
by: Zhao, Pu, et al.
Published: (2024)