Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhaskar, Uday, Bhattacharya, Rishabh, Patel, Avinash, Khoche, Sarthak, Kulkarni, Praveen Anil, Manwani, Naresh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915614918967296
author Bhaskar, Uday
Bhattacharya, Rishabh
Patel, Avinash
Khoche, Sarthak
Kulkarni, Praveen Anil
Manwani, Naresh
author_facet Bhaskar, Uday
Bhattacharya, Rishabh
Patel, Avinash
Khoche, Sarthak
Kulkarni, Praveen Anil
Manwani, Naresh
contents Foundation models, especially vision-language models (VLMs), offer compelling zero-shot object detection for applications like autonomous driving, a domain where manual labelling is prohibitively expensive. However, their detection latency and tendency to hallucinate predictions render them unsuitable for direct deployment. This work introduces a novel pipeline that addresses this challenge by leveraging VLMs to automatically generate pseudo-labels for training efficient, real-time object detectors. Our key innovation is a per-object co-teaching-based training strategy that mitigates the inherent noise in VLM-generated labels. The proposed per-object coteaching approach filters noisy bounding boxes from training instead of filtering the entire image. Specifically, two YOLO models learn collaboratively, filtering out unreliable boxes from each mini-batch based on their peers' per-object loss values. Overall, our pipeline provides an efficient, robust, and scalable approach to train high-performance object detectors for autonomous driving, significantly reducing reliance on costly human annotation. Experimental results on the KITTI dataset demonstrate that our method outperforms a baseline YOLOv5m model, achieving a significant mAP@0.5 boost ($31.12\%$ to $46.61\%$) while maintaining real-time detection latency. Furthermore, we show that supplementing our pseudo-labelled data with a small fraction of ground truth labels ($10\%$) leads to further performance gains, reaching $57.97\%$ mAP@0.5 on the KITTI dataset. We observe similar performance improvements for the ACDC and BDD100k datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09955
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching
Bhaskar, Uday
Bhattacharya, Rishabh
Patel, Avinash
Khoche, Sarthak
Kulkarni, Praveen Anil
Manwani, Naresh
Computer Vision and Pattern Recognition
Foundation models, especially vision-language models (VLMs), offer compelling zero-shot object detection for applications like autonomous driving, a domain where manual labelling is prohibitively expensive. However, their detection latency and tendency to hallucinate predictions render them unsuitable for direct deployment. This work introduces a novel pipeline that addresses this challenge by leveraging VLMs to automatically generate pseudo-labels for training efficient, real-time object detectors. Our key innovation is a per-object co-teaching-based training strategy that mitigates the inherent noise in VLM-generated labels. The proposed per-object coteaching approach filters noisy bounding boxes from training instead of filtering the entire image. Specifically, two YOLO models learn collaboratively, filtering out unreliable boxes from each mini-batch based on their peers' per-object loss values. Overall, our pipeline provides an efficient, robust, and scalable approach to train high-performance object detectors for autonomous driving, significantly reducing reliance on costly human annotation. Experimental results on the KITTI dataset demonstrate that our method outperforms a baseline YOLOv5m model, achieving a significant mAP@0.5 boost ($31.12\%$ to $46.61\%$) while maintaining real-time detection latency. Furthermore, we show that supplementing our pseudo-labelled data with a small fraction of ground truth labels ($10\%$) leads to further performance gains, reaching $57.97\%$ mAP@0.5 on the KITTI dataset. We observe similar performance improvements for the ACDC and BDD100k datasets.
title Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.09955