A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Wenze, Hosseinzadeh, Mehdi, Dayoub, Feras |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025)
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025)
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
von: Garg, Sourav, et al.
Veröffentlicht: (2024)
von: Garg, Sourav, et al.
Veröffentlicht: (2024)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024)
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)
Efficient Heatmap-Guided 6-Dof Grasp Detection in Cluttered Scenes
von: Chen, Siang, et al.
Veröffentlicht: (2024)
von: Chen, Siang, et al.
Veröffentlicht: (2024)
DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
Embodied Domain Adaptation for Object Detection
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
Agentic Pipeline for Self-Synchronized Multiview Joint Angle Monitoring in Uncalibrated Environments
von: Yu, Juncheng, et al.
Veröffentlicht: (2026)
von: Yu, Juncheng, et al.
Veröffentlicht: (2026)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
von: Back, Seunghyeok, et al.
Veröffentlicht: (2025)
von: Back, Seunghyeok, et al.
Veröffentlicht: (2025)
ObjectReact: Learning Object-Relative Control for Visual Navigation
von: Garg, Sourav, et al.
Veröffentlicht: (2025)
von: Garg, Sourav, et al.
Veröffentlicht: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
RealDex: Towards Human-like Grasping for Robotic Dexterous Hand
von: Liu, Yumeng, et al.
Veröffentlicht: (2024)
von: Liu, Yumeng, et al.
Veröffentlicht: (2024)
SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
von: Mei, Yunpeng, et al.
Veröffentlicht: (2025)
von: Mei, Yunpeng, et al.
Veröffentlicht: (2025)
Bimanual Grasp Synthesis for Dexterous Robot Hands
von: Shao, Yanming, et al.
Veröffentlicht: (2024)
von: Shao, Yanming, et al.
Veröffentlicht: (2024)
Robot Instance Segmentation with Few Annotations for Grasping
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)
Grasp-and-Lift: Executable 3D Hand-Object Interaction Reconstruction via Physics-in-the-Loop Optimization
von: Choi, Byeonggyeol, et al.
Veröffentlicht: (2026)
von: Choi, Byeonggyeol, et al.
Veröffentlicht: (2026)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
Robotic Grasping of Harvested Tomato Trusses Using Vision and Online Learning
von: Bent, Luuk van den, et al.
Veröffentlicht: (2023)
von: Bent, Luuk van den, et al.
Veröffentlicht: (2023)
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
von: Tsagkas, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsagkas, Nikolaos, et al.
Veröffentlicht: (2024)
Grasping Partially Occluded Objects Using Autoencoder-Based Point Cloud Inpainting
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
von: Fu, Jiahui, et al.
Veröffentlicht: (2026)
von: Fu, Jiahui, et al.
Veröffentlicht: (2026)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
von: Chen, Xinyi, et al.
Veröffentlicht: (2025)
von: Chen, Xinyi, et al.
Veröffentlicht: (2025)
RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
von: Feng, ZhiYuan, et al.
Veröffentlicht: (2026)
von: Feng, ZhiYuan, et al.
Veröffentlicht: (2026)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
von: Wu, Pengyuan, et al.
Veröffentlicht: (2026)
von: Wu, Pengyuan, et al.
Veröffentlicht: (2026)
Object-Centric World Model for Language-Guided Manipulation
von: Jeong, Youngjoon, et al.
Veröffentlicht: (2025)
von: Jeong, Youngjoon, et al.
Veröffentlicht: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
von: Zhang, Kaidong, et al.
Veröffentlicht: (2025)
von: Zhang, Kaidong, et al.
Veröffentlicht: (2025)
CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning
von: Ang, Sining, et al.
Veröffentlicht: (2026)
von: Ang, Sining, et al.
Veröffentlicht: (2026)
GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
von: Jiang, Guangqi, et al.
Veröffentlicht: (2025)
von: Jiang, Guangqi, et al.
Veröffentlicht: (2025)
PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking
von: Bao, Jiacheng, et al.
Veröffentlicht: (2026)
von: Bao, Jiacheng, et al.
Veröffentlicht: (2026)
RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset
von: Wang, Yongzhong, et al.
Veröffentlicht: (2026)
von: Wang, Yongzhong, et al.
Veröffentlicht: (2026)
PhotoBot: Reference-Guided Interactive Photography via Natural Language
von: Limoyo, Oliver, et al.
Veröffentlicht: (2024)
von: Limoyo, Oliver, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2026) -
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025) -
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
von: Garg, Sourav, et al.
Veröffentlicht: (2024) -
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024) -
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)