GAgent: An Adaptive Rigid-Soft Gripping Agent with Vision Language Models for Complex Lighting Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhuowei, Zhang, Miao, Lin, Xiaotian, Yin, Meng, Lu, Shuai, Wang, Xueqian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Interactive Navigation in Environments with Traversable Obstacles Using Large Language and Vision-Language Models
by: Zhang, Zhen, et al.
Published: (2023)
by: Zhang, Zhen, et al.
Published: (2023)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
by: Zhou, Hui, et al.
Published: (2025)
by: Zhou, Hui, et al.
Published: (2025)
SoftMAC: Differentiable Soft Body Simulation with Forecast-based Contact Model and Two-way Coupling with Articulated Rigid Bodies and Clothes
by: Liu, Min, et al.
Published: (2023)
by: Liu, Min, et al.
Published: (2023)
VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
by: Wu, Yuze, et al.
Published: (2025)
by: Wu, Yuze, et al.
Published: (2025)
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
by: Wang, Qiuyue, et al.
Published: (2026)
by: Wang, Qiuyue, et al.
Published: (2026)
Mastering Contact-rich Tasks by Combining Soft and Rigid Robotics with Imitation Learning
by: Montero, Mariano Ramírez, et al.
Published: (2024)
by: Montero, Mariano Ramírez, et al.
Published: (2024)
Vision Language Models Can Parse Floor Plan Maps
by: DeFazio, David, et al.
Published: (2024)
by: DeFazio, David, et al.
Published: (2024)
FalconWing: An Ultra-Light Indoor Fixed-Wing UAV Platform for Vision-Based Autonomy
by: Miao, Yan, et al.
Published: (2025)
by: Miao, Yan, et al.
Published: (2025)
A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking
by: Zhang, Yuelin, et al.
Published: (2026)
by: Zhang, Yuelin, et al.
Published: (2026)
Visual-tactile Fusion for Transparent Object Grasping in Complex Backgrounds
by: Li, Shoujie, et al.
Published: (2022)
by: Li, Shoujie, et al.
Published: (2022)
Adaptive Domain Modeling with Language Models: A Multi-Agent Approach to Task Planning
by: Babu, Harisankar, et al.
Published: (2025)
by: Babu, Harisankar, et al.
Published: (2025)
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
by: Su, Haochen, et al.
Published: (2025)
by: Su, Haochen, et al.
Published: (2025)
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
by: Yin, Cheng, et al.
Published: (2025)
by: Yin, Cheng, et al.
Published: (2025)
Adaptive Capacity Allocation for Vision Language Action Fine-tuning
by: Kim, Donghoon, et al.
Published: (2026)
by: Kim, Donghoon, et al.
Published: (2026)
Reduced-Order Model-Guided Reinforcement Learning for Demonstration-Free Humanoid Locomotion
by: Liu, Shuai, et al.
Published: (2025)
by: Liu, Shuai, et al.
Published: (2025)
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
by: Zhai, Shaopeng, et al.
Published: (2025)
by: Zhai, Shaopeng, et al.
Published: (2025)
AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models
by: Jia, Tingzheng, et al.
Published: (2026)
by: Jia, Tingzheng, et al.
Published: (2026)
General-Purpose Aerial Intelligent Agents Empowered by Large Language Models
by: Zhao, Ji, et al.
Published: (2025)
by: Zhao, Ji, et al.
Published: (2025)
Test-Time Adaptation for Tactile-Vision-Language Models
by: Ye, Chuyang, et al.
Published: (2026)
by: Ye, Chuyang, et al.
Published: (2026)
IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
by: Liufu, Weijia, et al.
Published: (2026)
by: Liufu, Weijia, et al.
Published: (2026)
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
by: Su, Taiyi, et al.
Published: (2026)
by: Su, Taiyi, et al.
Published: (2026)
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
by: Bai, Qianqian, et al.
Published: (2025)
by: Bai, Qianqian, et al.
Published: (2025)
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Zero-shot Object Navigation with Vision-Language Models Reasoning
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
by: Wei, Ziyu, et al.
Published: (2026)
by: Wei, Ziyu, et al.
Published: (2026)
SayNav: Grounding Large Language Models for Dynamic Planning to Navigation in New Environments
by: Rajvanshi, Abhinav, et al.
Published: (2023)
by: Rajvanshi, Abhinav, et al.
Published: (2023)
MA-VLCM: A Vision Language Critic Model for Value Estimation of Policies in Multi-Agent Team Settings
by: Shaik, Shahil, et al.
Published: (2026)
by: Shaik, Shahil, et al.
Published: (2026)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)
by: Zhang, Yihao, et al.
Published: (2025)
Learning Adaptive Hydrodynamic Models Using Neural ODEs in Complex Conditions
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
by: Peng, Jierui, et al.
Published: (2025)
by: Peng, Jierui, et al.
Published: (2025)
Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
CrashAgent: Crash Scenario Generation via Multi-modal Reasoning
by: Li, Miao, et al.
Published: (2025)
by: Li, Miao, et al.
Published: (2025)
UAV-CodeAgents: Scalable UAV Mission Planning via Multi-Agent ReAct and Vision-Language Reasoning
by: Sautenkov, Oleg, et al.
Published: (2025)
by: Sautenkov, Oleg, et al.
Published: (2025)
Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
by: Wang, Taowen, et al.
Published: (2024)
by: Wang, Taowen, et al.
Published: (2024)
SmoothVLA: Aligning Vision-Language-Action Models with Physical Constraints via Intrinsic Smoothness Optimization
by: Li, Jiashun, et al.
Published: (2026)
by: Li, Jiashun, et al.
Published: (2026)
Embodied Instruction Following in Unknown Environments
by: Wu, Zhenyu, et al.
Published: (2024)
by: Wu, Zhenyu, et al.
Published: (2024)
Similar Items
-
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments
by: Li, Zhiyuan, et al.
Published: (2024) -
Interactive Navigation in Environments with Traversable Obstacles Using Large Language and Vision-Language Models
by: Zhang, Zhen, et al.
Published: (2023) -
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
by: Zhou, Hui, et al.
Published: (2025) -
SoftMAC: Differentiable Soft Body Simulation with Forecast-based Contact Model and Two-way Coupling with Articulated Rigid Bodies and Clothes
by: Liu, Min, et al.
Published: (2023) -
VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
by: Wu, Yuze, et al.
Published: (2025)