CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917173280112640 |
|---|---|
| author | Raoufi, Behnam Sharify, Hossein Ramezanee, Mohamad Mahdee Hajsadeghi, Khosrow Shouraki, Saeed Bagheri |
| author_facet | Raoufi, Behnam Sharify, Hossein Ramezanee, Mohamad Mahdee Hajsadeghi, Khosrow Shouraki, Saeed Bagheri |
| contents | Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Joint-Detect, a simple and detector-agnostic framework that integrates CLIP-style contrastive vision-language supervision through end-to-end joint training. A lightweight parallel head projects region or grid features into the CLIP embedding space and aligns them with learnable class-specific text embeddings via InfoNCE contrastive loss and an auxiliary cross-entropy term, while all standard detection losses are optimized simultaneously. The approach applies seamlessly to both two-stage and one-stage architectures. We validate it on Pascal VOC 2007+2012 using Faster R-CNN and on the large-scale MS COCO 2017 benchmark using modern YOLO detectors (YOLOv11), achieving consistent and substantial improvements while preserving real-time inference speed. Extensive experiments and ablations demonstrate that joint optimization with learnable text embeddings markedly enhances closed-set detection performance across diverse architectures and datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_22969 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision Raoufi, Behnam Sharify, Hossein Ramezanee, Mohamad Mahdee Hajsadeghi, Khosrow Shouraki, Saeed Bagheri Computer Vision and Pattern Recognition I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10; I.4.8 Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Joint-Detect, a simple and detector-agnostic framework that integrates CLIP-style contrastive vision-language supervision through end-to-end joint training. A lightweight parallel head projects region or grid features into the CLIP embedding space and aligns them with learnable class-specific text embeddings via InfoNCE contrastive loss and an auxiliary cross-entropy term, while all standard detection losses are optimized simultaneously. The approach applies seamlessly to both two-stage and one-stage architectures. We validate it on Pascal VOC 2007+2012 using Faster R-CNN and on the large-scale MS COCO 2017 benchmark using modern YOLO detectors (YOLOv11), achieving consistent and substantial improvements while preserving real-time inference speed. Extensive experiments and ablations demonstrate that joint optimization with learnable text embeddings markedly enhances closed-set detection performance across diverse architectures and datasets. |
| title | CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision |
| topic | Computer Vision and Pattern Recognition I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10, I.4.8 I.2.10; I.4.8 |
| url | https://arxiv.org/abs/2512.22969 |