Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yitong, Yao, Wenhao, Meng, Lingchen, Wu, Sihong, Wu, Zuxuan, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
by: Chen, Yitong, et al.
Published: (2025)
by: Chen, Yitong, et al.
Published: (2025)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
by: Meng, Lingchen, et al.
Published: (2024)
by: Meng, Lingchen, et al.
Published: (2024)
CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
by: Sun, Zhichao, et al.
Published: (2025)
by: Sun, Zhichao, et al.
Published: (2025)
Enhanced Object Detection: A Study on Vast Vocabulary Object Detection Track for V3Det Challenge 2024
by: Wu, Peixi, et al.
Published: (2024)
by: Wu, Peixi, et al.
Published: (2024)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
by: Chen, Yitong, et al.
Published: (2026)
by: Chen, Yitong, et al.
Published: (2026)
SEGIC: Unleashing the Emergent Correspondence for In-Context Segmentation
by: Meng, Lingchen, et al.
Published: (2023)
by: Meng, Lingchen, et al.
Published: (2023)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
by: Peng, Wujian, et al.
Published: (2024)
by: Peng, Wujian, et al.
Published: (2024)
VastTrack: Vast Category Visual Object Tracking
by: Peng, Liang, et al.
Published: (2024)
by: Peng, Liang, et al.
Published: (2024)
Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Multi-Prompt Alignment for Multi-Source Unsupervised Domain Adaptation
by: Chen, Haoran, et al.
Published: (2022)
by: Chen, Haoran, et al.
Published: (2022)
FOCUS: Towards Universal Foreground Segmentation
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
by: Wang, Junke, et al.
Published: (2023)
by: Wang, Junke, et al.
Published: (2023)
Learning Accurate Segmentation Purely from Self-Supervision
by: You, Zuyao, et al.
Published: (2026)
by: You, Zuyao, et al.
Published: (2026)
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation
by: Chen, Haoran, et al.
Published: (2025)
by: Chen, Haoran, et al.
Published: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
by: Zhou, Ziwei, et al.
Published: (2025)
by: Zhou, Ziwei, et al.
Published: (2025)
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
by: Yao, Jianhang, et al.
Published: (2025)
by: Yao, Jianhang, et al.
Published: (2025)
VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction
by: Lin, Jiaqi, et al.
Published: (2024)
by: Lin, Jiaqi, et al.
Published: (2024)
DiffusionAD: Norm-guided One-step Denoising Diffusion for Anomaly Detection
by: Zhang, Hui, et al.
Published: (2023)
by: Zhang, Hui, et al.
Published: (2023)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
by: Wang, Junke, et al.
Published: (2025)
by: Wang, Junke, et al.
Published: (2025)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection
by: Wang, Wanqi, et al.
Published: (2026)
by: Wang, Wanqi, et al.
Published: (2026)
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
by: Liu, Zhuohan, et al.
Published: (2026)
by: Liu, Zhuohan, et al.
Published: (2026)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
by: Zou, Zichen, et al.
Published: (2026)
by: Zou, Zichen, et al.
Published: (2026)
GeoGS3D: Single-view 3D Reconstruction via Geometric-aware Diffusion Model and Gaussian Splatting
by: Feng, Qijun, et al.
Published: (2024)
by: Feng, Qijun, et al.
Published: (2024)
State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
by: Zhou, Jiaying, et al.
Published: (2025)
by: Zhou, Jiaying, et al.
Published: (2025)
Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection
by: Weng, Zhenhai, et al.
Published: (2025)
by: Weng, Zhenhai, et al.
Published: (2025)
Multi-Modal Decouple and Recouple Network for Robust 3D Object Detection
by: Ding, Rui, et al.
Published: (2026)
by: Ding, Rui, et al.
Published: (2026)
Pro-AD: Learning Comprehensive Prototypes with Prototype-based Constraint for Multi-class Unsupervised Anomaly Detection
by: Zhou, Ziqing, et al.
Published: (2025)
by: Zhou, Ziqing, et al.
Published: (2025)
BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection
by: Li, Zhenxin, et al.
Published: (2023)
by: Li, Zhenxin, et al.
Published: (2023)
Evaluating the Performance of Open-Vocabulary Object Detection in Low-quality Image
by: Wu, Po-Chih
Published: (2025)
by: Wu, Po-Chih
Published: (2025)
VoxelNextFusion: A Simple, Unified and Effective Voxel Fusion Framework for Multi-Modal 3D Object Detection
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object Detection
by: Zhang, Gang, et al.
Published: (2024)
by: Zhang, Gang, et al.
Published: (2024)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
by: Chen, Haoran, et al.
Published: (2024)
by: Chen, Haoran, et al.
Published: (2024)
PromptFusion: Decoupling Stability and Plasticity for Continual Learning
by: Chen, Haoran, et al.
Published: (2023)
by: Chen, Haoran, et al.
Published: (2023)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
by: Weng, Zejia, et al.
Published: (2024)
by: Weng, Zejia, et al.
Published: (2024)
Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
by: Zhao, Youjun, et al.
Published: (2025)
by: Zhao, Youjun, et al.
Published: (2025)
Similar Items
-
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
by: Wang, Jiaqi, et al.
Published: (2024) -
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
by: Chen, Yitong, et al.
Published: (2025) -
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
by: Meng, Lingchen, et al.
Published: (2024) -
CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
by: Sun, Zhichao, et al.
Published: (2025) -
Enhanced Object Detection: A Study on Vast Vocabulary Object Detection Track for V3Det Challenge 2024
by: Wu, Peixi, et al.
Published: (2024)