CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Raoufi, Behnam, Sharify, Hossein, Ramezanee, Mohamad Mahdee, Hajsadeghi, Khosrow, Shouraki, Saeed Bagheri
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917173280112640
author Raoufi, Behnam
Sharify, Hossein
Ramezanee, Mohamad Mahdee
Hajsadeghi, Khosrow
Shouraki, Saeed Bagheri
author_facet Raoufi, Behnam
Sharify, Hossein
Ramezanee, Mohamad Mahdee
Hajsadeghi, Khosrow
Shouraki, Saeed Bagheri
contents Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Joint-Detect, a simple and detector-agnostic framework that integrates CLIP-style contrastive vision-language supervision through end-to-end joint training. A lightweight parallel head projects region or grid features into the CLIP embedding space and aligns them with learnable class-specific text embeddings via InfoNCE contrastive loss and an auxiliary cross-entropy term, while all standard detection losses are optimized simultaneously. The approach applies seamlessly to both two-stage and one-stage architectures. We validate it on Pascal VOC 2007+2012 using Faster R-CNN and on the large-scale MS COCO 2017 benchmark using modern YOLO detectors (YOLOv11), achieving consistent and substantial improvements while preserving real-time inference speed. Extensive experiments and ablations demonstrate that joint optimization with learnable text embeddings markedly enhances closed-set detection performance across diverse architectures and datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22969
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
Raoufi, Behnam
Sharify, Hossein
Ramezanee, Mohamad Mahdee
Hajsadeghi, Khosrow
Shouraki, Saeed Bagheri
Computer Vision and Pattern Recognition
I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8
I.2.10; I.4.8
Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Joint-Detect, a simple and detector-agnostic framework that integrates CLIP-style contrastive vision-language supervision through end-to-end joint training. A lightweight parallel head projects region or grid features into the CLIP embedding space and aligns them with learnable class-specific text embeddings via InfoNCE contrastive loss and an auxiliary cross-entropy term, while all standard detection losses are optimized simultaneously. The approach applies seamlessly to both two-stage and one-stage architectures. We validate it on Pascal VOC 2007+2012 using Faster R-CNN and on the large-scale MS COCO 2017 benchmark using modern YOLO detectors (YOLOv11), achieving consistent and substantial improvements while preserving real-time inference speed. Extensive experiments and ablations demonstrate that joint optimization with learnable text embeddings markedly enhances closed-set detection performance across diverse architectures and datasets.
title CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
topic Computer Vision and Pattern Recognition
I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8
I.2.10; I.4.8
url https://arxiv.org/abs/2512.22969