PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
Fuente:
arXiv
Saved in:
| Main Authors: | Dao, Alan, Vu, Dinh Bach, Anh, Tuan Le Duc, Huy, Bui Quang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
Georegistration Improvements Using ArUco Markers as Ground Control Points for Fast Deployment of 3D Map Reconstruction
by: Anh Quang Nguyen, et al.
Published: (2025)
by: Anh Quang Nguyen, et al.
Published: (2025)
D2S: Representing sparse descriptors and 3D coordinates for camera relocalization
by: Bui, Bach-Thuan, et al.
Published: (2023)
by: Bui, Bach-Thuan, et al.
Published: (2023)
Design of a Bio-Inspired Miniature Submarine for Low-Cost Water Quality Monitoring
by: Vu, Quang Huy, et al.
Published: (2026)
by: Vu, Quang Huy, et al.
Published: (2026)
T2Nav Algebraic Topology Aware Temporal Graph Memory and Loop Detection for ZeroShot Visual Navigation
by: D., Quang-Anh N., et al.
Published: (2026)
by: D., Quang-Anh N., et al.
Published: (2026)
A Heuristic Motion Planning Algorithm for a Mobile Robot With Nonholonomic Constraints
by: Duc Thien Tran, et al.
Published: (2024)
by: Duc Thien Tran, et al.
Published: (2024)
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
by: Dao, Alan, et al.
Published: (2024)
by: Dao, Alan, et al.
Published: (2024)
Vision-Guided Targeted Grasping and Vibration for Robotic Pollination in Controlled Environments
by: Jeong, Jaehwan, et al.
Published: (2025)
by: Jeong, Jaehwan, et al.
Published: (2025)
Jan-nano Technical Report
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
EFHQ: Multi-purpose ExtremePose-Face-HQ dataset
by: Dao, Trung Tuan, et al.
Published: (2023)
by: Dao, Trung Tuan, et al.
Published: (2023)
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
by: Vuong, An Dinh, et al.
Published: (2025)
by: Vuong, An Dinh, et al.
Published: (2025)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
by: Cao-Dinh, Duc, et al.
Published: (2025)
by: Cao-Dinh, Duc, et al.
Published: (2025)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
by: Long, Keke, et al.
Published: (2024)
by: Long, Keke, et al.
Published: (2024)
Enhancing Depth Image Estimation for Underwater Robots by Combining Image Processing and Machine Learning
by: Nguyen, Quang Truong, et al.
Published: (2024)
by: Nguyen, Quang Truong, et al.
Published: (2024)
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
Evaluation of an Autonomous Surface Robot Equipped with a Transformable Mobility Mechanism for Efficient Mobility Control
by: Fujii, Yasuyuki, et al.
Published: (2025)
by: Fujii, Yasuyuki, et al.
Published: (2025)
Autonomous Block Assembly for Boom Cranes with Passive Joint Dynamics: Integrated Vision MPC Control
by: Ebmer, Gerald, et al.
Published: (2026)
by: Ebmer, Gerald, et al.
Published: (2026)
X-ray Fluoroscopy Guided Localization and Steering of Medical Microrobots through Virtual Enhancement
by: Alabay, Husnu Halid, et al.
Published: (2024)
by: Alabay, Husnu Halid, et al.
Published: (2024)
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
by: Vo, Khoa, et al.
Published: (2026)
by: Vo, Khoa, et al.
Published: (2026)
Improved 3D Point-Line Mapping Regression for Camera Relocalization
by: Bui, Bach-Thuan, et al.
Published: (2025)
by: Bui, Bach-Thuan, et al.
Published: (2025)
Dr-PoGO: Direct Radar Pose-Graph Optimization
by: Gentil, Cedric Le, et al.
Published: (2026)
by: Gentil, Cedric Le, et al.
Published: (2026)
A comparison of extended object tracking with multi-modal sensors in indoor environment
by: Shuai, Jiangtao, et al.
Published: (2024)
by: Shuai, Jiangtao, et al.
Published: (2024)
HumanoidVLM: Vision-Language-Guided Impedance Control for Contact-Rich Humanoid Manipulation
by: Mahmoud, Yara, et al.
Published: (2026)
by: Mahmoud, Yara, et al.
Published: (2026)
Simulator Adaptation for Sim-to-Real Learning of Legged Locomotion via Proprioceptive Distribution Matching
by: Dao, Jeremy, et al.
Published: (2026)
by: Dao, Jeremy, et al.
Published: (2026)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Language-Driven Closed-Loop Grasping with Model-Predictive Trajectory Replanning
by: Nguyen, Huy Hoang, et al.
Published: (2024)
by: Nguyen, Huy Hoang, et al.
Published: (2024)
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
by: Li, Yuquan, et al.
Published: (2026)
by: Li, Yuquan, et al.
Published: (2026)
Grasping, Part Identification, and Pose Refinement in One Shot with a Tactile Gripper
by: Lim, Joyce Xin-Yan, et al.
Published: (2023)
by: Lim, Joyce Xin-Yan, et al.
Published: (2023)
RGBTrack: Fast, Robust Depth-Free 6D Pose Estimation and Tracking
by: Guo, Teng, et al.
Published: (2025)
by: Guo, Teng, et al.
Published: (2025)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation
by: Le, Huy, et al.
Published: (2024)
by: Le, Huy, et al.
Published: (2024)
Agentic AI Meets Edge Computing in Autonomous UAV Swarms
by: Nguyen, Thuan Minh, et al.
Published: (2026)
by: Nguyen, Thuan Minh, et al.
Published: (2026)
Bio-Inspired Hybrid Map: Spatial Implicit Local Frames and Topological Map for Mobile Cobot Navigation
by: Dang, Tuan, et al.
Published: (2025)
by: Dang, Tuan, et al.
Published: (2025)
Global Existence of Solutions to Semilinear σ‐Evolution Equations With Different Damping Types in Lq Framework
by: Dinh Van Duong, et al.
Published: (2025)
by: Dinh Van Duong, et al.
Published: (2025)
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
by: Nguyen, Huy Hoang, et al.
Published: (2024)
by: Nguyen, Huy Hoang, et al.
Published: (2024)
Active Tactile Exploration for Rigid Body Pose and Shape Estimation
by: Gordon, Ethan K., et al.
Published: (2025)
by: Gordon, Ethan K., et al.
Published: (2025)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2025)
by: Van Vo, Tuan, et al.
Published: (2025)
Overcoming Dynamic Environments: A Hybrid Approach to Motion Planning for Manipulators
by: Ngo, Ho Minh Quang, et al.
Published: (2025)
by: Ngo, Ho Minh Quang, et al.
Published: (2025)
Similar Items
-
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning
by: Dao, Alan, et al.
Published: (2025) -
Georegistration Improvements Using ArUco Markers as Ground Control Points for Fast Deployment of 3D Map Reconstruction
by: Anh Quang Nguyen, et al.
Published: (2025) -
D2S: Representing sparse descriptors and 3D coordinates for camera relocalization
by: Bui, Bach-Thuan, et al.
Published: (2023) -
Design of a Bio-Inspired Miniature Submarine for Low-Cost Water Quality Monitoring
by: Vu, Quang Huy, et al.
Published: (2026) -
T2Nav Algebraic Topology Aware Temporal Graph Memory and Loop Detection for ZeroShot Visual Navigation
by: D., Quang-Anh N., et al.
Published: (2026)