LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yang, Zhao, Shiyu, Chen, Yuxiao, Wang, Zhenting, Jin, Can, Metaxas, Dimitris N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915502898544640
author Zhou, Yang
Zhao, Shiyu
Chen, Yuxiao
Wang, Zhenting
Jin, Can
Metaxas, Dimitris N.
author_facet Zhou, Yang
Zhao, Shiyu
Chen, Yuxiao
Wang, Zhenting
Jin, Can
Metaxas, Dimitris N.
contents Large foundation models trained on large-scale vision-language data can boost Open-Vocabulary Object Detection (OVD) via synthetic training data, yet the hand-crafted pipelines often introduce bias and overfit to specific prompts. We sidestep this issue by directly fusing hidden states from Large Language Models (LLMs) into detectors-an avenue surprisingly under-explored. This paper presents a systematic method to enhance visual grounding by utilizing decoder layers of the LLM of an MLLM. We introduce a zero-initialized cross-attention adapter to enable efficient knowledge fusion from LLMs to object detectors, a new approach called LED (LLM Enhanced Open-Vocabulary Object Detection). We find that intermediate LLM layers already encode rich spatial semantics; adapting only the early layers yields most of the gain. With Swin-T as the vision encoder, Qwen2-0.5B + LED lifts GroundingDINO by 3.82 % on OmniLabel at just 8.7 % extra GFLOPs, and a larger vision backbone pushes the improvement to 6.22 %. Extensive ablations on adapter variants, LLM scales and fusion depths further corroborate our design.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
Zhou, Yang
Zhao, Shiyu
Chen, Yuxiao
Wang, Zhenting
Jin, Can
Metaxas, Dimitris N.
Computer Vision and Pattern Recognition
Artificial Intelligence
Large foundation models trained on large-scale vision-language data can boost Open-Vocabulary Object Detection (OVD) via synthetic training data, yet the hand-crafted pipelines often introduce bias and overfit to specific prompts. We sidestep this issue by directly fusing hidden states from Large Language Models (LLMs) into detectors-an avenue surprisingly under-explored. This paper presents a systematic method to enhance visual grounding by utilizing decoder layers of the LLM of an MLLM. We introduce a zero-initialized cross-attention adapter to enable efficient knowledge fusion from LLMs to object detectors, a new approach called LED (LLM Enhanced Open-Vocabulary Object Detection). We find that intermediate LLM layers already encode rich spatial semantics; adapting only the early layers yields most of the gain. With Swin-T as the vision encoder, Qwen2-0.5B + LED lifts GroundingDINO by 3.82 % on OmniLabel at just 8.7 % extra GFLOPs, and a larger vision backbone pushes the improvement to 6.22 %. Extensive ablations on adapter variants, LLM scales and fusion depths further corroborate our design.
title LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.13794