Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
Fuente:
arXiv
Guardado en:
| Autores principales: | Mei, Guofeng, Riz, Luigi, Wang, Yiming, Poiesi, Fabio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
por: Mei, Guofeng, et al.
Publicado: (2026)
por: Mei, Guofeng, et al.
Publicado: (2026)
Geometrically-driven Aggregation for Zero-shot 3D Point Cloud Understanding
por: Mei, Guofeng, et al.
Publicado: (2023)
por: Mei, Guofeng, et al.
Publicado: (2023)
PerLA: Perceptive 3D Language Assistant
por: Mei, Guofeng, et al.
Publicado: (2024)
por: Mei, Guofeng, et al.
Publicado: (2024)
Multimodal Fusion SLAM with Fourier Attention
por: Zhou, Youjie, et al.
Publicado: (2025)
por: Zhou, Youjie, et al.
Publicado: (2025)
Novel class discovery meets foundation models for 3D semantic segmentation
por: Riz, Luigi, et al.
Publicado: (2023)
por: Riz, Luigi, et al.
Publicado: (2023)
Free-form language-based robotic reasoning and grasping
por: Jiao, Runyu, et al.
Publicado: (2025)
por: Jiao, Runyu, et al.
Publicado: (2025)
Obstruction reasoning for robotic grasping
por: Jiao, Runyu, et al.
Publicado: (2025)
por: Jiao, Runyu, et al.
Publicado: (2025)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
por: Kim, Juno, et al.
Publicado: (2025)
por: Kim, Juno, et al.
Publicado: (2025)
Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
por: Mei, Guofeng, et al.
Publicado: (2025)
por: Mei, Guofeng, et al.
Publicado: (2025)
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
por: Wang, Chenhao, et al.
Publicado: (2026)
por: Wang, Chenhao, et al.
Publicado: (2026)
MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment
por: Li, Bingyu, et al.
Publicado: (2025)
por: Li, Bingyu, et al.
Publicado: (2025)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
por: Barsellotti, Luca, et al.
Publicado: (2024)
por: Barsellotti, Luca, et al.
Publicado: (2024)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
por: Zhu, Yuanbing, et al.
Publicado: (2024)
por: Zhu, Yuanbing, et al.
Publicado: (2024)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
por: Luo, Jiayun, et al.
Publicado: (2023)
por: Luo, Jiayun, et al.
Publicado: (2023)
Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
por: Nguyen, Binh Long, et al.
Publicado: (2026)
por: Nguyen, Binh Long, et al.
Publicado: (2026)
Training-Free Unsupervised Prompt for Vision-Language Models
por: Long, Sifan, et al.
Publicado: (2024)
por: Long, Sifan, et al.
Publicado: (2024)
XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation
por: Wang, Ziyi, et al.
Publicado: (2024)
por: Wang, Ziyi, et al.
Publicado: (2024)
GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
por: Tao, Xujing, et al.
Publicado: (2026)
por: Tao, Xujing, et al.
Publicado: (2026)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
por: Vu, Tuan-Anh, et al.
Publicado: (2023)
por: Vu, Tuan-Anh, et al.
Publicado: (2023)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
por: Xie, Jiahao, et al.
Publicado: (2023)
por: Xie, Jiahao, et al.
Publicado: (2023)
Light-weight Retinal Layer Segmentation with Global Reasoning
por: He, Xiang, et al.
Publicado: (2024)
por: He, Xiang, et al.
Publicado: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
por: Guo, Xuechen, et al.
Publicado: (2024)
por: Guo, Xuechen, et al.
Publicado: (2024)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
por: Nguyen, Phuc D. A., et al.
Publicado: (2024)
por: Nguyen, Phuc D. A., et al.
Publicado: (2024)
Deep Learning-Based 3D Instance and Semantic Segmentation: A Review
por: Yasir, Siddiqui Muhammad, et al.
Publicado: (2024)
por: Yasir, Siddiqui Muhammad, et al.
Publicado: (2024)
Foveated Instance Segmentation
por: Zeng, Hongyi, et al.
Publicado: (2025)
por: Zeng, Hongyi, et al.
Publicado: (2025)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
por: Jin, Yizhang, et al.
Publicado: (2024)
por: Jin, Yizhang, et al.
Publicado: (2024)
Language-Guided Instance-Aware Domain-Adaptive Panoptic Segmentation
por: Mansour, Elham Amin, et al.
Publicado: (2024)
por: Mansour, Elham Amin, et al.
Publicado: (2024)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
por: Li, Jinlong, et al.
Publicado: (2025)
por: Li, Jinlong, et al.
Publicado: (2025)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
por: Dutta, Saikat, et al.
Publicado: (2026)
por: Dutta, Saikat, et al.
Publicado: (2026)
Unsupervised Instance Segmentation with Superpixels
por: Hoang, Cuong Manh
Publicado: (2025)
por: Hoang, Cuong Manh
Publicado: (2025)
GeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video
por: Sharma, Arun
Publicado: (2026)
por: Sharma, Arun
Publicado: (2026)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
por: Xue, Feng, et al.
Publicado: (2025)
por: Xue, Feng, et al.
Publicado: (2025)
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
por: Chen, Yanhui, et al.
Publicado: (2026)
por: Chen, Yanhui, et al.
Publicado: (2026)
RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
por: Li, Ke, et al.
Publicado: (2025)
por: Li, Ke, et al.
Publicado: (2025)
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision
por: Wang, Zhaoqing, et al.
Publicado: (2024)
por: Wang, Zhaoqing, et al.
Publicado: (2024)
FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
por: Pang, Kaicheng, et al.
Publicado: (2025)
por: Pang, Kaicheng, et al.
Publicado: (2025)
Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
por: Huang, Qiming, et al.
Publicado: (2025)
por: Huang, Qiming, et al.
Publicado: (2025)
Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
por: Goel, Atharv, et al.
Publicado: (2025)
por: Goel, Atharv, et al.
Publicado: (2025)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
por: Zeng, Haoxi, et al.
Publicado: (2026)
por: Zeng, Haoxi, et al.
Publicado: (2026)
GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence
por: Xiao, Qinfeng, et al.
Publicado: (2026)
por: Xiao, Qinfeng, et al.
Publicado: (2026)
Ejemplares similares
-
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
por: Mei, Guofeng, et al.
Publicado: (2026) -
Geometrically-driven Aggregation for Zero-shot 3D Point Cloud Understanding
por: Mei, Guofeng, et al.
Publicado: (2023) -
PerLA: Perceptive 3D Language Assistant
por: Mei, Guofeng, et al.
Publicado: (2024) -
Multimodal Fusion SLAM with Fourier Attention
por: Zhou, Youjie, et al.
Publicado: (2025) -
Novel class discovery meets foundation models for 3D semantic segmentation
por: Riz, Luigi, et al.
Publicado: (2023)