LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Shenghao, Yang, Qize, Mo, Qijie, Yan, Junkai, Wei, Xihan, Meng, Jingke, Xie, Xiaohua, Zheng, Wei-Shi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024)
by: Fu, Shenghao, et al.
Published: (2024)
Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
by: Mo, Qijie, et al.
Published: (2024)
by: Mo, Qijie, et al.
Published: (2024)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
by: Yan, Junkai, et al.
Published: (2024)
by: Yan, Junkai, et al.
Published: (2024)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)
by: Peng, Yi-Xing, et al.
Published: (2025)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
by: Tian, Huilin, et al.
Published: (2024)
by: Tian, Huilin, et al.
Published: (2024)
PRET: Planning with Directed Fidelity Trajectory for Vision and Language Navigation
by: Lu, Renjie, et al.
Published: (2024)
by: Lu, Renjie, et al.
Published: (2024)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Domain Adaptation for Large-Vocabulary Object Detectors
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts
by: Chhipa, Prakash Chandra, et al.
Published: (2024)
by: Chhipa, Prakash Chandra, et al.
Published: (2024)
The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection
by: Wan, Yazhe, et al.
Published: (2026)
by: Wan, Yazhe, et al.
Published: (2026)
Structured Spatial Reasoning with Open Vocabulary Object Detectors
by: Nejatishahidin, Negar, et al.
Published: (2024)
by: Nejatishahidin, Negar, et al.
Published: (2024)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
by: Garcia-Fernandez, Pablo, et al.
Published: (2025)
by: Garcia-Fernandez, Pablo, et al.
Published: (2025)
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-world Object Detector
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
by: Ruis, Frank, et al.
Published: (2025)
by: Ruis, Frank, et al.
Published: (2025)
ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation
by: Zeng, Ling-An, et al.
Published: (2025)
by: Zeng, Ling-An, et al.
Published: (2025)
Exploring Open-Vocabulary Object Recognition in Images using CLIP
by: Chen, Wei Yu, et al.
Published: (2026)
by: Chen, Wei Yu, et al.
Published: (2026)
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
by: Li, Badi, et al.
Published: (2025)
by: Li, Badi, et al.
Published: (2025)
Unsupervised Open-Vocabulary Object Localization in Videos
by: Fan, Ke, et al.
Published: (2023)
by: Fan, Ke, et al.
Published: (2023)
OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding
by: Yang, Dianyi, et al.
Published: (2025)
by: Yang, Dianyi, et al.
Published: (2025)
Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning
by: Wei, Guoyizhe, et al.
Published: (2026)
by: Wei, Guoyizhe, et al.
Published: (2026)
Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
by: Mütze, Annika, et al.
Published: (2025)
by: Mütze, Annika, et al.
Published: (2025)
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
by: Raj, Ankita, et al.
Published: (2025)
by: Raj, Ankita, et al.
Published: (2025)
Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization
by: Du, Jia-Run, et al.
Published: (2024)
by: Du, Jia-Run, et al.
Published: (2024)
Rethinking Few-shot Class-incremental Learning: Learning from Yourself
by: Tang, Yu-Ming, et al.
Published: (2024)
by: Tang, Yu-Ming, et al.
Published: (2024)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
by: Zhao, Kai, et al.
Published: (2025)
by: Zhao, Kai, et al.
Published: (2025)
VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking
by: Qian, Zekun, et al.
Published: (2024)
by: Qian, Zekun, et al.
Published: (2024)
State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
by: Zhou, Jiaying, et al.
Published: (2025)
by: Zhou, Jiaying, et al.
Published: (2025)
Open-Vocabulary Object Detection via Language Hierarchy
by: Huang, Jiaxing, et al.
Published: (2024)
by: Huang, Jiaxing, et al.
Published: (2024)
Vision-Language Models are Strong Noisy Label Detectors
by: Wei, Tong, et al.
Published: (2024)
by: Wei, Tong, et al.
Published: (2024)
osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning
by: Xie, Fujing, et al.
Published: (2025)
by: Xie, Fujing, et al.
Published: (2025)
Similar Items
-
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025) -
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024) -
Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
by: Mo, Qijie, et al.
Published: (2024) -
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025) -
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)