Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Pätzold, Bastian, Nogga, Jan, Behnke, Sven |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
by: Choy, Chris, et al.
Published: (2026)
by: Choy, Chris, et al.
Published: (2026)
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
by: Cao, Helin, et al.
Published: (2025)
by: Cao, Helin, et al.
Published: (2025)
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
by: Quenzel, Jan, et al.
Published: (2024)
by: Quenzel, Jan, et al.
Published: (2024)
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
by: Villar-Corrales, Angel, et al.
Published: (2024)
by: Villar-Corrales, Angel, et al.
Published: (2024)
PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2024)
by: Yu, Xuan, et al.
Published: (2024)
Person Segmentation and Action Classification for Multi-Channel Hemisphere Field of View LiDAR Sensors
by: Seliunina, Svetlana, et al.
Published: (2024)
by: Seliunina, Svetlana, et al.
Published: (2024)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models
by: Cao, Helin, et al.
Published: (2024)
by: Cao, Helin, et al.
Published: (2024)
Leverage Cross-Attention for End-to-End Open-Vocabulary Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
by: Lu, Yangxiao, et al.
Published: (2024)
by: Lu, Yangxiao, et al.
Published: (2024)
Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking
by: Ishaq, Ayesha, et al.
Published: (2024)
by: Ishaq, Ayesha, et al.
Published: (2024)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-Net
by: Cao, Helin, et al.
Published: (2024)
by: Cao, Helin, et al.
Published: (2024)
Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation
by: Tutevych, Vitalii, et al.
Published: (2026)
by: Tutevych, Vitalii, et al.
Published: (2026)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
Adapting Segment Anything Model for Unseen Object Instance Segmentation
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
Open-Set 3D Semantic Instance Maps for Vision Language Navigation -- O3D-SIM
by: Nanwani, Laksh, et al.
Published: (2024)
by: Nanwani, Laksh, et al.
Published: (2024)
Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models
by: Nishimura, Takayuki, et al.
Published: (2024)
by: Nishimura, Takayuki, et al.
Published: (2024)
FM-Fusion: Instance-aware Semantic Mapping Boosted by Vision-Language Foundation Models
by: Liu, Chuhao, et al.
Published: (2024)
by: Liu, Chuhao, et al.
Published: (2024)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
by: Cao, Helin, et al.
Published: (2025)
by: Cao, Helin, et al.
Published: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
by: Zhang, Qiyao, et al.
Published: (2026)
by: Zhang, Qiyao, et al.
Published: (2026)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
by: Piekenbrinck, Jens, et al.
Published: (2025)
by: Piekenbrinck, Jens, et al.
Published: (2025)
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
by: Zhang, Qifan, et al.
Published: (2026)
by: Zhang, Qifan, et al.
Published: (2026)
Are Open-Vocabulary Models Ready for Detection of MEP Elements on Construction Sites
by: Abdalwhab, Abdalwhab, et al.
Published: (2025)
by: Abdalwhab, Abdalwhab, et al.
Published: (2025)
LOVON: Legged Open-Vocabulary Object Navigator
by: Peng, Daojie, et al.
Published: (2025)
by: Peng, Daojie, et al.
Published: (2025)
Open-Vocabulary Online Semantic Mapping for SLAM
by: Martins, Tomas Berriel, et al.
Published: (2024)
by: Martins, Tomas Berriel, et al.
Published: (2024)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling
by: Qiu, Xiaowen, et al.
Published: (2025)
by: Qiu, Xiaowen, et al.
Published: (2025)
Towards Open-World Grasping with Large Vision-Language Models
by: Tziafas, Georgios, et al.
Published: (2024)
by: Tziafas, Georgios, et al.
Published: (2024)
OVerSeeC: Open-Vocabulary Costmap Generation from Satellite Images and Natural Language
by: Rana, Rwik, et al.
Published: (2026)
by: Rana, Rwik, et al.
Published: (2026)
SENSE: Stereo OpEN Vocabulary SEmantic Segmentation
by: Campagnolo, Thomas, et al.
Published: (2026)
by: Campagnolo, Thomas, et al.
Published: (2026)
WildOS: Open-Vocabulary Object Search in the Wild
by: Shah, Hardik, et al.
Published: (2026)
by: Shah, Hardik, et al.
Published: (2026)
OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
by: Zhou, Zhishan, et al.
Published: (2025)
by: Zhou, Zhishan, et al.
Published: (2025)
Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussian
by: Chahe, Amirhosein, et al.
Published: (2024)
by: Chahe, Amirhosein, et al.
Published: (2024)
DualMap: Online Open-Vocabulary Semantic Mapping for Natural Language Navigation in Dynamic Changing Scenes
by: Jiang, Jiajun, et al.
Published: (2025)
by: Jiang, Jiajun, et al.
Published: (2025)
OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation
by: Jiang, Haochen, et al.
Published: (2024)
by: Jiang, Haochen, et al.
Published: (2024)
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
by: Wu, Yanmin, et al.
Published: (2024)
by: Wu, Yanmin, et al.
Published: (2024)
Similar Items
-
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
by: Choy, Chris, et al.
Published: (2026) -
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
by: Cao, Helin, et al.
Published: (2025) -
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
by: Quenzel, Jan, et al.
Published: (2024) -
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
by: Villar-Corrales, Angel, et al.
Published: (2024) -
PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2024)