High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Quan-Sheng, Li, Yunheng, Zhou, Daquan, Li, Guanbin, Hou, Qibin, Cheng, Ming-Ming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
by: Zhang, Shi-Chen, et al.
Published: (2025)
by: Zhang, Shi-Chen, et al.
Published: (2025)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
by: Zeng, Quan-Sheng, et al.
Published: (2025)
by: Zeng, Quan-Sheng, et al.
Published: (2025)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
by: Li, Yunheng, et al.
Published: (2025)
by: Li, Yunheng, et al.
Published: (2025)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024)
by: Li, Xuanyi, et al.
Published: (2024)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024)
by: Li, Yongkang, et al.
Published: (2024)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
Open-Vocabulary Segmentation with Semantic-Assisted Calibration
by: Liu, Yong, et al.
Published: (2023)
by: Liu, Yong, et al.
Published: (2023)
PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling
by: Li, Zhong-Yu, et al.
Published: (2024)
by: Li, Zhong-Yu, et al.
Published: (2024)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023)
by: Yin, Bowen, et al.
Published: (2023)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
by: Zhang, Xuying, et al.
Published: (2025)
by: Zhang, Xuying, et al.
Published: (2025)
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
by: Li, Yu-Jhe, et al.
Published: (2024)
by: Li, Yu-Jhe, et al.
Published: (2024)
Contrastive Masked Autoencoders are Stronger Vision Learners
by: Huang, Zhicheng, et al.
Published: (2022)
by: Huang, Zhicheng, et al.
Published: (2022)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
by: Wang, Kuo, et al.
Published: (2024)
by: Wang, Kuo, et al.
Published: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
by: Zhou, Yupeng, et al.
Published: (2023)
by: Zhou, Yupeng, et al.
Published: (2023)
OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation
by: Zhao, Ganlong, et al.
Published: (2024)
by: Zhao, Ganlong, et al.
Published: (2024)
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2026)
by: Li, Jiaming, et al.
Published: (2026)
SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
by: Xiong, Xinyu, et al.
Published: (2025)
by: Xiong, Xinyu, et al.
Published: (2025)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026)
by: Zhou, Yupeng, et al.
Published: (2026)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
by: Zhang, Xuying, et al.
Published: (2024)
by: Zhang, Xuying, et al.
Published: (2024)
Zone Evaluation: Revealing Spatial Bias in Object Detection
by: Zheng, Zhaohui, et al.
Published: (2023)
by: Zheng, Zhaohui, et al.
Published: (2023)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter
by: Wang, Jinglong, et al.
Published: (2023)
by: Wang, Jinglong, et al.
Published: (2023)
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting
by: Li, Di, et al.
Published: (2025)
by: Li, Di, et al.
Published: (2025)
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024)
by: Kim, Chanyoung, et al.
Published: (2024)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
by: Zhang, Dengke, et al.
Published: (2024)
by: Zhang, Dengke, et al.
Published: (2024)
MaskClustering: View Consensus based Mask Graph Clustering for Open-Vocabulary 3D Instance Segmentation
by: Yan, Mi, et al.
Published: (2024)
by: Yan, Mi, et al.
Published: (2024)
MedSeg-R: Medical Image Segmentation with Clinical Reasoning
by: Shao, Hao, et al.
Published: (2025)
by: Shao, Hao, et al.
Published: (2025)
Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
by: Nguyen, Phuc D. A., et al.
Published: (2023)
by: Nguyen, Phuc D. A., et al.
Published: (2023)
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
by: Yuan, Haobo, et al.
Published: (2024)
by: Yuan, Haobo, et al.
Published: (2024)
Similar Items
-
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
by: Li, Yunheng, et al.
Published: (2024) -
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024) -
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
by: Zhang, Shi-Chen, et al.
Published: (2025) -
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
by: Zeng, Quan-Sheng, et al.
Published: (2025) -
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
by: Li, Yunheng, et al.
Published: (2025)