C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Siheng, Li, Zhengdao, Li, Yanshu, Xiao, Canran, Zhan, Haibo, Yao, Zhengtao, Zhang, Xuzhi, Kang, Jiale, Li, Linshan, Liu, Weiming, Dong, Zhikang, Shen, Jifeng, Dong, Junhao, Sun, Qiang, Koniusz, Piotr
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909959058358272
author Wang, Siheng
Li, Zhengdao
Li, Yanshu
Xiao, Canran
Zhan, Haibo
Yao, Zhengtao
Zhang, Xuzhi
Kang, Jiale
Li, Linshan
Liu, Weiming
Dong, Zhikang
Shen, Jifeng
Dong, Junhao
Sun, Qiang
Koniusz, Piotr
author_facet Wang, Siheng
Li, Zhengdao
Li, Yanshu
Xiao, Canran
Zhan, Haibo
Yao, Zhengtao
Zhang, Xuzhi
Kang, Jiale
Li, Linshan
Liu, Weiming
Dong, Zhikang
Shen, Jifeng
Dong, Junhao
Sun, Qiang
Koniusz, Piotr
contents Object detection has advanced significantly in the closed-set setting, but real-world deployment remains limited by two challenges: poor generalization to unseen categories and insufficient robustness under adverse conditions. Prior research has explored these issues separately: visible-infrared detection improves robustness but lacks generalization, while open-world detection leverages vision-language alignment strategy for category diversity but struggles under extreme environments. This trade-off leaves robustness and diversity difficult to achieve simultaneously. To mitigate these issues, we propose \textbf{C3-OWD}, a curriculum cross-modal contrastive learning framework that unifies both strengths. Stage~1 enhances robustness by pretraining with RGBT data, while Stage~2 improves generalization via vision-language alignment. To prevent catastrophic forgetting between two stages, we introduce an Exponential Moving Average (EMA) mechanism that theoretically guarantees preservation of pre-stage performance with bounded parameter lag and function consistency. Experiments on FLIR, OV-COCO, and OV-LVIS demonstrate the effectiveness of our approach: C3-OWD achieves $80.1$ AP$^{50}$ on FLIR, $48.6$ AP$^{50}_{\text{Novel}}$ on OV-COCO, and $35.7$ mAP$_r$ on OV-LVIS, establishing competitive performance across both robustness and diversity evaluations. Code available at: https://github.com/justin-herry/C3-OWD.git.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
Wang, Siheng
Li, Zhengdao
Li, Yanshu
Xiao, Canran
Zhan, Haibo
Yao, Zhengtao
Zhang, Xuzhi
Kang, Jiale
Li, Linshan
Liu, Weiming
Dong, Zhikang
Shen, Jifeng
Dong, Junhao
Sun, Qiang
Koniusz, Piotr
Computer Vision and Pattern Recognition
Object detection has advanced significantly in the closed-set setting, but real-world deployment remains limited by two challenges: poor generalization to unseen categories and insufficient robustness under adverse conditions. Prior research has explored these issues separately: visible-infrared detection improves robustness but lacks generalization, while open-world detection leverages vision-language alignment strategy for category diversity but struggles under extreme environments. This trade-off leaves robustness and diversity difficult to achieve simultaneously. To mitigate these issues, we propose \textbf{C3-OWD}, a curriculum cross-modal contrastive learning framework that unifies both strengths. Stage~1 enhances robustness by pretraining with RGBT data, while Stage~2 improves generalization via vision-language alignment. To prevent catastrophic forgetting between two stages, we introduce an Exponential Moving Average (EMA) mechanism that theoretically guarantees preservation of pre-stage performance with bounded parameter lag and function consistency. Experiments on FLIR, OV-COCO, and OV-LVIS demonstrate the effectiveness of our approach: C3-OWD achieves $80.1$ AP$^{50}$ on FLIR, $48.6$ AP$^{50}_{\text{Novel}}$ on OV-COCO, and $35.7$ mAP$_r$ on OV-LVIS, establishing competitive performance across both robustness and diversity evaluations. Code available at: https://github.com/justin-herry/C3-OWD.git.
title C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.23316