Zero-Shot Dual-Path Integration Framework for Open-Vocabulary 3D Instance Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ton, Tri, Hong, Ji Woo, Eom, SooHwan, Shim, Jun Yeop, Kim, Junyeong, Yoo, Chang D. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
di: Eom, SooHwan, et al.
Pubblicazione: (2026)
di: Eom, SooHwan, et al.
Pubblicazione: (2026)
TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
di: Ton, Tri, et al.
Pubblicazione: (2025)
di: Ton, Tri, et al.
Pubblicazione: (2025)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
di: Hong, Ji Woo, et al.
Pubblicazione: (2026)
di: Hong, Ji Woo, et al.
Pubblicazione: (2026)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
di: Eom, SooHwan, et al.
Pubblicazione: (2024)
di: Eom, SooHwan, et al.
Pubblicazione: (2024)
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
A Simple Framework for Open-Vocabulary Zero-Shot Segmentation
di: Stegmüller, Thomas, et al.
Pubblicazione: (2024)
di: Stegmüller, Thomas, et al.
Pubblicazione: (2024)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
di: Kim, Juno, et al.
Pubblicazione: (2025)
di: Kim, Juno, et al.
Pubblicazione: (2025)
MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation
di: Pham, Trung X., et al.
Pubblicazione: (2024)
di: Pham, Trung X., et al.
Pubblicazione: (2024)
ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On
di: Hong, Ji Woo, et al.
Pubblicazione: (2025)
di: Hong, Ji Woo, et al.
Pubblicazione: (2025)
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
di: Lebailly, Tim, et al.
Pubblicazione: (2025)
di: Lebailly, Tim, et al.
Pubblicazione: (2025)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
di: Yoon, Eunseop, et al.
Pubblicazione: (2025)
di: Yoon, Eunseop, et al.
Pubblicazione: (2025)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
di: Unmesh, Asim, et al.
Pubblicazione: (2026)
di: Unmesh, Asim, et al.
Pubblicazione: (2026)
Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning
di: Nguyen, Thanh, et al.
Pubblicazione: (2026)
di: Nguyen, Thanh, et al.
Pubblicazione: (2026)
E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization
di: Pham, Trung X., et al.
Pubblicazione: (2025)
di: Pham, Trung X., et al.
Pubblicazione: (2025)
DOZE: A Dataset for Open-Vocabulary Zero-Shot Object Navigation in Dynamic Environments
di: Ma, Ji, et al.
Pubblicazione: (2024)
di: Ma, Ji, et al.
Pubblicazione: (2024)
Towards Real-Time Open-Vocabulary Video Instance Segmentation
di: Yan, Bin, et al.
Pubblicazione: (2024)
di: Yan, Bin, et al.
Pubblicazione: (2024)
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
di: Fang, Hao, et al.
Pubblicazione: (2024)
di: Fang, Hao, et al.
Pubblicazione: (2024)
Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
di: Choi, Jiho, et al.
Pubblicazione: (2024)
di: Choi, Jiho, et al.
Pubblicazione: (2024)
YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation
di: Sapkota, Ranjan, et al.
Pubblicazione: (2026)
di: Sapkota, Ranjan, et al.
Pubblicazione: (2026)
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
di: Nguyen, Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Thanh, et al.
Pubblicazione: (2024)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
di: Pätzold, Bastian, et al.
Pubblicazione: (2025)
di: Pätzold, Bastian, et al.
Pubblicazione: (2025)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
di: Zhu, Wenqi, et al.
Pubblicazione: (2024)
di: Zhu, Wenqi, et al.
Pubblicazione: (2024)
Zero-Shot Scene Change Detection
di: Cho, Kyusik, et al.
Pubblicazione: (2024)
di: Cho, Kyusik, et al.
Pubblicazione: (2024)
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
di: Yoon, Eunseop, et al.
Pubblicazione: (2024)
di: Yoon, Eunseop, et al.
Pubblicazione: (2024)
A Zero-Shot Open-Vocabulary Pipeline for Dialogue Understanding
di: Safa, Abdulfattah, et al.
Pubblicazione: (2024)
di: Safa, Abdulfattah, et al.
Pubblicazione: (2024)
RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
di: Alama, Omar, et al.
Pubblicazione: (2025)
di: Alama, Omar, et al.
Pubblicazione: (2025)
Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation
di: Ulmer, Maximilian, et al.
Pubblicazione: (2025)
di: Ulmer, Maximilian, et al.
Pubblicazione: (2025)
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
di: Cheng, Zesen, et al.
Pubblicazione: (2024)
di: Cheng, Zesen, et al.
Pubblicazione: (2024)
SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation
di: Huang, Shiqi, et al.
Pubblicazione: (2025)
di: Huang, Shiqi, et al.
Pubblicazione: (2025)
MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment
di: Li, Bingyu, et al.
Pubblicazione: (2025)
di: Li, Bingyu, et al.
Pubblicazione: (2025)
Spatiotemporal Patterns of Fish Diversity in the Waters Around the Five West Sea Islands of South Korea: Integrating Bottom Trawl and Environmental DNA (eDNA) Methods.
di: Yoo, Young-Ji, et al.
Pubblicazione: (2025)
di: Yoo, Young-Ji, et al.
Pubblicazione: (2025)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
di: Chu, Wen-Hsuan, et al.
Pubblicazione: (2023)
di: Chu, Wen-Hsuan, et al.
Pubblicazione: (2023)
ZS-VCOS: Zero-Shot Video Camouflaged Object Segmentation By Optical Flow and Open Vocabulary Object Detection
di: Guo, Wenqi, et al.
Pubblicazione: (2025)
di: Guo, Wenqi, et al.
Pubblicazione: (2025)
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation
di: Ahn, Jinwoo, et al.
Pubblicazione: (2024)
di: Ahn, Jinwoo, et al.
Pubblicazione: (2024)
OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
di: Zhou, Zhishan, et al.
Pubblicazione: (2025)
di: Zhou, Zhishan, et al.
Pubblicazione: (2025)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
di: Boudjoghra, Mohamed El Amine, et al.
Pubblicazione: (2024)
di: Boudjoghra, Mohamed El Amine, et al.
Pubblicazione: (2024)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
di: Piekenbrinck, Jens, et al.
Pubblicazione: (2025)
di: Piekenbrinck, Jens, et al.
Pubblicazione: (2025)
Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
di: Nguyen, Phuc D. A., et al.
Pubblicazione: (2023)
di: Nguyen, Phuc D. A., et al.
Pubblicazione: (2023)
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
di: Choi, Jiho, et al.
Pubblicazione: (2025)
di: Choi, Jiho, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
di: Eom, SooHwan, et al.
Pubblicazione: (2026) -
TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
di: Ton, Tri, et al.
Pubblicazione: (2025) -
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
di: Hong, Ji Woo, et al.
Pubblicazione: (2026) -
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
di: Eom, SooHwan, et al.
Pubblicazione: (2024) -
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)