VLM6D: VLM based 6Dof Pose Estimation based on RGB-D Images
Fuente:
arXiv
Saved in:
| Main Authors: | Sarowar, Md Selim, Kim, Sungho |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
by: Sarowar, Md Selim, et al.
Published: (2025)
by: Sarowar, Md Selim, et al.
Published: (2025)
Explainable Parkinsons Disease Gait Recognition Using Multimodal RGB-D Fusion and Large Language Models
by: Alnaasan, Manar, et al.
Published: (2025)
by: Alnaasan, Manar, et al.
Published: (2025)
RDPN6D: Residual-based Dense Point-wise Network for 6Dof Object Pose Estimation Based on RGB-D Images
by: Hong, Zong-Wei, et al.
Published: (2024)
by: Hong, Zong-Wei, et al.
Published: (2024)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
by: Sarowar, Md Selim, et al.
Published: (2026)
by: Sarowar, Md Selim, et al.
Published: (2026)
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
by: Qian, Chen, et al.
Published: (2025)
by: Qian, Chen, et al.
Published: (2025)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
EfficientPose 6D: Scalable and Efficient 6D Object Pose Estimation
by: Fang, Zixuan, et al.
Published: (2025)
by: Fang, Zixuan, et al.
Published: (2025)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
Domain Adaptation of VLM for Soccer Video Understanding
by: Jiang, Tiancheng, et al.
Published: (2025)
by: Jiang, Tiancheng, et al.
Published: (2025)
Efficient Heatmap-Guided 6-Dof Grasp Detection in Cluttered Scenes
by: Chen, Siang, et al.
Published: (2024)
by: Chen, Siang, et al.
Published: (2024)
Extending 6D Object Pose Estimators for Stereo Vision
by: Pöllabauer, Thomas, et al.
Published: (2024)
by: Pöllabauer, Thomas, et al.
Published: (2024)
Any6D: Model-free 6D Pose Estimation of Novel Objects
by: Lee, Taeyeop, et al.
Published: (2025)
by: Lee, Taeyeop, et al.
Published: (2025)
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
by: Lee, Yuxiao, et al.
Published: (2025)
by: Lee, Yuxiao, et al.
Published: (2025)
Uncertainty Quantification with Deep Ensembles for 6D Object Pose Estimation
by: Wursthorn, Kira, et al.
Published: (2024)
by: Wursthorn, Kira, et al.
Published: (2024)
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
by: Wen, Bowen, et al.
Published: (2023)
by: Wen, Bowen, et al.
Published: (2023)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
by: Singh, Aditya Kumar, et al.
Published: (2026)
by: Singh, Aditya Kumar, et al.
Published: (2026)
WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
Dynamic VLM-Guided Negative Prompting for Diffusion Models
by: Chang, Hoyeon, et al.
Published: (2025)
by: Chang, Hoyeon, et al.
Published: (2025)
GS2Pose: Two-stage 6D Object Pose Estimation Guided by Gaussian Splatting
by: Mei, Jilan, et al.
Published: (2024)
by: Mei, Jilan, et al.
Published: (2024)
Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
by: Singhania, Aditi, et al.
Published: (2025)
by: Singhania, Aditi, et al.
Published: (2025)
Improving 6D Object Pose Estimation of metallic Household and Industry Objects
by: Pöllabauer, Thomas, et al.
Published: (2025)
by: Pöllabauer, Thomas, et al.
Published: (2025)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
by: Baik, Sangwon, et al.
Published: (2026)
by: Baik, Sangwon, et al.
Published: (2026)
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
by: Sun, Fan-Yun, et al.
Published: (2024)
by: Sun, Fan-Yun, et al.
Published: (2024)
DecomPose: Disentangling Cross-Category Optimization Contention for Category-Level 6D Object Pose Estimation
by: Gao, Yifan, et al.
Published: (2026)
by: Gao, Yifan, et al.
Published: (2026)
FAST GDRNPP: Improving the Speed of State-of-the-Art 6D Object Pose Estimation
by: Pöllabauer, Thomas, et al.
Published: (2024)
by: Pöllabauer, Thomas, et al.
Published: (2024)
Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
by: Kim, Junsu, et al.
Published: (2025)
by: Kim, Junsu, et al.
Published: (2025)
Box6D : Zero-shot Category-level 6D Pose Estimation of Warehouse Boxes
by: Ma, Yintao, et al.
Published: (2025)
by: Ma, Yintao, et al.
Published: (2025)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
by: Wei, Tong, et al.
Published: (2025)
by: Wei, Tong, et al.
Published: (2025)
Trust but Verify: Programmatic VLM Evaluation in the Wild
by: Prabhu, Viraj, et al.
Published: (2024)
by: Prabhu, Viraj, et al.
Published: (2024)
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
by: Zhang, Mengchen, et al.
Published: (2024)
by: Zhang, Mengchen, et al.
Published: (2024)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
by: Fu, Yongjie, et al.
Published: (2024)
by: Fu, Yongjie, et al.
Published: (2024)
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
by: Xie, Roy, et al.
Published: (2026)
by: Xie, Roy, et al.
Published: (2026)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
SmolVLM: Redefining small and efficient multimodal models
by: Marafioti, Andrés, et al.
Published: (2025)
by: Marafioti, Andrés, et al.
Published: (2025)
PushupBench: Your VLM is not good at counting pushups
by: Li, Shengzhi, et al.
Published: (2026)
by: Li, Shengzhi, et al.
Published: (2026)
Similar Items
-
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
by: Sarowar, Md Selim, et al.
Published: (2025) -
Explainable Parkinsons Disease Gait Recognition Using Multimodal RGB-D Fusion and Large Language Models
by: Alnaasan, Manar, et al.
Published: (2025) -
RDPN6D: Residual-based Dense Point-wise Network for 6Dof Object Pose Estimation Based on RGB-D Images
by: Hong, Zong-Wei, et al.
Published: (2024) -
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
by: Sarowar, Md Selim, et al.
Published: (2026) -
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
by: Qian, Chen, et al.
Published: (2025)