A3VLM: Actionable Articulation-Aware Vision Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Siyuan, Chang, Haonan, Liu, Yuhan, Zhu, Yimeng, Dong, Hao, Gao, Peng, Boularias, Abdeslam, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
Scaling Manipulation Learning with Visual Kinematic Chain Prediction
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
Autoregressive Action Sequence Learning for Robotic Manipulation
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
von: Yu, Qiaojun, et al.
Veröffentlicht: (2024)
von: Yu, Qiaojun, et al.
Veröffentlicht: (2024)
Bellman Diffusion Models
von: Schramm, Liam, et al.
Veröffentlicht: (2024)
von: Schramm, Liam, et al.
Veröffentlicht: (2024)
Bounding Distributional Shifts in World Modeling through Novelty Detection
von: Jing, Eric, et al.
Veröffentlicht: (2025)
von: Jing, Eric, et al.
Veröffentlicht: (2025)
DAP: Diffusion-based Affordance Prediction for Multi-modality Storage
von: Chang, Haonan, et al.
Veröffentlicht: (2024)
von: Chang, Haonan, et al.
Veröffentlicht: (2024)
Provably Efficient Long-Horizon Exploration in Monte Carlo Tree Search through State Occupancy Regularization
von: Schramm, Liam, et al.
Veröffentlicht: (2024)
von: Schramm, Liam, et al.
Veröffentlicht: (2024)
One-Shot Imitation Learning with Invariance Matching for Robotic Manipulation
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
von: Li, Xin, et al.
Veröffentlicht: (2024)
von: Li, Xin, et al.
Veröffentlicht: (2024)
KARL: Kalman-Filter Assisted Reinforcement Learner for Dynamic Object Tracking and Grasping
von: Boyalakuntla, Kowndinya, et al.
Veröffentlicht: (2025)
von: Boyalakuntla, Kowndinya, et al.
Veröffentlicht: (2025)
LGMCTS: Language-Guided Monte-Carlo Tree Search for Executable Semantic Object Rearrangement
von: Chang, Haonan, et al.
Veröffentlicht: (2023)
von: Chang, Haonan, et al.
Veröffentlicht: (2023)
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
PROBE: Proprioceptive Obstacle Detection and Estimation while Navigating in Clutter
von: Ramesh, Dhruv Metha, et al.
Veröffentlicht: (2025)
von: Ramesh, Dhruv Metha, et al.
Veröffentlicht: (2025)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
von: Long, Keke, et al.
Veröffentlicht: (2024)
von: Long, Keke, et al.
Veröffentlicht: (2024)
Learning Visual Feature-Based World Models via Residual Latent Action
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
Integrating Model-based Control and RL for Sim2Real Transfer of Tight Insertion Policies
von: Marougkas, Isidoros, et al.
Veröffentlicht: (2025)
von: Marougkas, Isidoros, et al.
Veröffentlicht: (2025)
SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects
von: Geng, Haoran, et al.
Veröffentlicht: (2023)
von: Geng, Haoran, et al.
Veröffentlicht: (2023)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving
von: Qu, Yansong, et al.
Veröffentlicht: (2025)
von: Qu, Yansong, et al.
Veröffentlicht: (2025)
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
von: Liu, Jiaxing, et al.
Veröffentlicht: (2026)
von: Liu, Jiaxing, et al.
Veröffentlicht: (2026)
Detect Everything with Few Examples
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models
von: Song, Daeun, et al.
Veröffentlicht: (2024)
von: Song, Daeun, et al.
Veröffentlicht: (2024)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
von: Kim, Taewhan, et al.
Veröffentlicht: (2024)
von: Kim, Taewhan, et al.
Veröffentlicht: (2024)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
von: Wu, Yuhan, et al.
Veröffentlicht: (2025)
von: Wu, Yuhan, et al.
Veröffentlicht: (2025)
GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects
von: Yu, Qiaojun, et al.
Veröffentlicht: (2023)
von: Yu, Qiaojun, et al.
Veröffentlicht: (2023)
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise
von: Ling, Suhan, et al.
Veröffentlicht: (2024)
von: Ling, Suhan, et al.
Veröffentlicht: (2024)
CoINS: Counterfactual Interactive Navigation via Skill-Aware VLM
von: Zhou, Kangjie, et al.
Veröffentlicht: (2026)
von: Zhou, Kangjie, et al.
Veröffentlicht: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
VLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing
von: Tang, Guoqin, et al.
Veröffentlicht: (2026)
von: Tang, Guoqin, et al.
Veröffentlicht: (2026)
Pandora: Articulated 3D Scene Graphs from Egocentric Vision
von: Yu, Alan, et al.
Veröffentlicht: (2026)
von: Yu, Alan, et al.
Veröffentlicht: (2026)
AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
von: Mishra, Sarthak, et al.
Veröffentlicht: (2025)
von: Mishra, Sarthak, et al.
Veröffentlicht: (2025)
SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
von: He, Yumeng, et al.
Veröffentlicht: (2025)
von: He, Yumeng, et al.
Veröffentlicht: (2025)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
von: Lou, Zhichen, et al.
Veröffentlicht: (2025)
von: Lou, Zhichen, et al.
Veröffentlicht: (2025)
Vi-TacMan: Articulated Object Manipulation via Vision and Touch
von: Cui, Leiyao, et al.
Veröffentlicht: (2025)
von: Cui, Leiyao, et al.
Veröffentlicht: (2025)
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
von: Zhu, Shaoting, et al.
Veröffentlicht: (2024)
von: Zhu, Shaoting, et al.
Veröffentlicht: (2024)
ReplanVLM: Replanning Robotic Tasks with Visual Language Models
von: Mei, Aoran, et al.
Veröffentlicht: (2024)
von: Mei, Aoran, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies
von: Liu, Yuhan, et al.
Veröffentlicht: (2025) -
Scaling Manipulation Learning with Visual Kinematic Chain Prediction
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024) -
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025) -
Autoregressive Action Sequence Learning for Robotic Manipulation
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024) -
UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
von: Yu, Qiaojun, et al.
Veröffentlicht: (2024)