Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yajima, Masaru, Ota, Kei, Kanezaki, Asako, Kawakami, Rei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Touch2Insert: Zero-Shot Peg Insertion by Touching Intersections of Peg and Hole
by: Yajima, Masaru, et al.
Published: (2026)
by: Yajima, Masaru, et al.
Published: (2026)
FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
by: Wu, Kuanting, et al.
Published: (2025)
by: Wu, Kuanting, et al.
Published: (2025)
Embodied Navigation with Auxiliary Task of Action Description Prediction
by: Kondoh, Haru, et al.
Published: (2025)
by: Kondoh, Haru, et al.
Published: (2025)
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
by: Sun, Leyuan, et al.
Published: (2024)
by: Sun, Leyuan, et al.
Published: (2024)
OP-Align: Object-level and Part-level Alignment for Self-supervised Category-level Articulated Object Pose Estimation
by: Che, Yuchen, et al.
Published: (2024)
by: Che, Yuchen, et al.
Published: (2024)
UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
by: Jiang, Zhaodong, et al.
Published: (2025)
by: Jiang, Zhaodong, et al.
Published: (2025)
COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
by: Che, Yuchen, et al.
Published: (2026)
by: Che, Yuchen, et al.
Published: (2026)
Zero-shot Degree of Ill-posedness Estimation for Active Small Object Change Detection
by: Takeda, Koji, et al.
Published: (2024)
by: Takeda, Koji, et al.
Published: (2024)
From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
by: Pulli, Tessa, et al.
Published: (2024)
by: Pulli, Tessa, et al.
Published: (2024)
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
by: Liu, Katherine, et al.
Published: (2025)
by: Liu, Katherine, et al.
Published: (2025)
SpyroPose: SE(3) Pyramids for Object Pose Distribution Estimation
by: Haugaard, Rasmus Laurvig, et al.
Published: (2023)
by: Haugaard, Rasmus Laurvig, et al.
Published: (2023)
SurgPose: Generalisable Surgical Instrument Pose Estimation using Zero-Shot Learning and Stereo Vision
by: Rai, Utsav, et al.
Published: (2025)
by: Rai, Utsav, et al.
Published: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)
by: Chen, Kehan, et al.
Published: (2024)
Color-Pair Guided Robust Zero-Shot 6D Pose Estimation and Tracking of Cluttered Objects on Edge Devices
by: Yang, Xingjian, et al.
Published: (2025)
by: Yang, Xingjian, et al.
Published: (2025)
SE(3)-PoseFlow: Estimating 6D Pose Distributions for Uncertainty-Aware Robotic Manipulation
by: Jin, Yufeng, et al.
Published: (2025)
by: Jin, Yufeng, et al.
Published: (2025)
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
MonoSE(3)-Diffusion: A Monocular SE(3) Diffusion Framework for Robust Camera-to-Robot Pose Estimation
by: Zhu, Kangjian, et al.
Published: (2025)
by: Zhu, Kangjian, et al.
Published: (2025)
Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
by: Qiao, Yanyuan, et al.
Published: (2024)
by: Qiao, Yanyuan, et al.
Published: (2024)
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
by: Zheng, Wanrong, et al.
Published: (2026)
by: Zheng, Wanrong, et al.
Published: (2026)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
HIPPo: Harnessing Image-to-3D Priors for Model-free Zero-shot 6D Pose Estimation
by: Liu, Yibo, et al.
Published: (2025)
by: Liu, Yibo, et al.
Published: (2025)
Zero-Splat TeleAssist: A Zero-Shot Pose Estimation Framework for Semantic Teleoperation
by: Dokania, Srijan, et al.
Published: (2025)
by: Dokania, Srijan, et al.
Published: (2025)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
by: Kapelyukh, Ivan, et al.
Published: (2023)
by: Kapelyukh, Ivan, et al.
Published: (2023)
High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions
by: Cao, Yongpeng, et al.
Published: (2026)
by: Cao, Yongpeng, et al.
Published: (2026)
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
by: Lyu, Kailin, et al.
Published: (2026)
by: Lyu, Kailin, et al.
Published: (2026)
Correspondence-Free Pose Estimation with Patterns: A Unified Approach for Multi-Dimensional Vision
by: Quan, Quan, et al.
Published: (2025)
by: Quan, Quan, et al.
Published: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement
by: Newman, Benjamin A., et al.
Published: (2024)
by: Newman, Benjamin A., et al.
Published: (2024)
Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models
by: Khalil, Aws, et al.
Published: (2026)
by: Khalil, Aws, et al.
Published: (2026)
SPADE: Sparsity Adaptive Depth Estimator for Zero-Shot, Real-Time, Monocular Depth Estimation in Underwater Environments
by: Zhang, Hongjie, et al.
Published: (2025)
by: Zhang, Hongjie, et al.
Published: (2025)
FoundPose: Unseen Object Pose Estimation with Foundation Features
by: Örnek, Evin Pınar, et al.
Published: (2023)
by: Örnek, Evin Pınar, et al.
Published: (2023)
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models
by: Huynh, Cuong, et al.
Published: (2026)
by: Huynh, Cuong, et al.
Published: (2026)
Modality Selection and Skill Segmentation via Cross-Modality Attention
by: Jiang, Jiawei, et al.
Published: (2025)
by: Jiang, Jiawei, et al.
Published: (2025)
DriveVA: Video Action Models are Zero-Shot Drivers
by: Liu, Mengmeng, et al.
Published: (2026)
by: Liu, Mengmeng, et al.
Published: (2026)
ZeroSCD: Zero-Shot Street Scene Change Detection
by: Kannan, Shyam Sundar, et al.
Published: (2024)
by: Kannan, Shyam Sundar, et al.
Published: (2024)
ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
by: Ren, Huan, et al.
Published: (2026)
by: Ren, Huan, et al.
Published: (2026)
Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models
by: Ünal, Berkehan, et al.
Published: (2026)
by: Ünal, Berkehan, et al.
Published: (2026)
Similar Items
-
Touch2Insert: Zero-Shot Peg Insertion by Touching Intersections of Peg and Hole
by: Yajima, Masaru, et al.
Published: (2026) -
FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
by: Wu, Kuanting, et al.
Published: (2025) -
Embodied Navigation with Auxiliary Task of Action Description Prediction
by: Kondoh, Haru, et al.
Published: (2025) -
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
by: Sun, Leyuan, et al.
Published: (2024) -
OP-Align: Object-level and Part-level Alignment for Self-supervised Category-level Articulated Object Pose Estimation
by: Che, Yuchen, et al.
Published: (2024)