CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Chao, Yin, Jianqin, Zhang, Ren
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912320913932288
author Qi, Chao
Yin, Jianqin
Zhang, Ren
author_facet Qi, Chao
Yin, Jianqin
Zhang, Ren
contents Image-point class incremental learning helps the 3D-points-vision robots continually learn category knowledge from 2D images, improving their perceptual capability in dynamic environments. However, some incremental learning methods address unimodal forgetting but fail in cross-modal cases, while others handle modal differences within training/testing datasets but assume no modal gaps between them. We first explore this cross-modal task, proposing a benchmark CMIP-CIL and relieving the cross-modal catastrophic forgetting problem. It employs masked point clouds and rendered multi-view images within a contrastive learning framework in pre-training, empowering the vision model with the generalizations of image-point correspondence. In the incremental stage, by freezing the backbone and promoting object representations close to their respective prototypes, the model effectively retains and generalizes knowledge across previously seen categories while continuing to learn new ones. We conduct comprehensive experiments on the benchmark datasets. Experiments prove that our method achieves state-of-the-art results, outperforming the baseline methods by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08422
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning
Qi, Chao
Yin, Jianqin
Zhang, Ren
Computer Vision and Pattern Recognition
Image-point class incremental learning helps the 3D-points-vision robots continually learn category knowledge from 2D images, improving their perceptual capability in dynamic environments. However, some incremental learning methods address unimodal forgetting but fail in cross-modal cases, while others handle modal differences within training/testing datasets but assume no modal gaps between them. We first explore this cross-modal task, proposing a benchmark CMIP-CIL and relieving the cross-modal catastrophic forgetting problem. It employs masked point clouds and rendered multi-view images within a contrastive learning framework in pre-training, empowering the vision model with the generalizations of image-point correspondence. In the incremental stage, by freezing the backbone and promoting object representations close to their respective prototypes, the model effectively retains and generalizes knowledge across previously seen categories while continuing to learn new ones. We conduct comprehensive experiments on the benchmark datasets. Experiments prove that our method achieves state-of-the-art results, outperforming the baseline methods by a large margin.
title CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.08422