Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ye, Hua, Ding, Hang, Chen, Siyuan, Jiang, Yiyang, Zhang, Changyuan, Zhang, Xuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917197996097536
author Ye, Hua
Ding, Hang
Chen, Siyuan
Jiang, Yiyang
Zhang, Changyuan
Zhang, Xuan
author_facet Ye, Hua
Ding, Hang
Chen, Siyuan
Jiang, Yiyang
Zhang, Changyuan
Zhang, Xuan
contents Most multimodal models treat every negative pair alike, ignoring the ambiguous negatives that differ from the positive by only a small detail. We propose Boundary-Aware Curriculum with Local Attention (BACL), a lightweight add-on that turns these borderline cases into a curriculum signal. A Boundary-aware Negative Sampler gradually raises difficulty, while a Contrastive Local Attention loss highlights where the mismatch occurs. The two modules are fully differentiable and work with any off-the-shelf dual encoder. Theory predicts a fast O(1/n) error rate; practice shows up to +32% R@1 over CLIP and new SOTA on four large-scale benchmarks, all without extra labels.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
Ye, Hua
Ding, Hang
Chen, Siyuan
Jiang, Yiyang
Zhang, Changyuan
Zhang, Xuan
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.6; I.2.10; I.2.7
Most multimodal models treat every negative pair alike, ignoring the ambiguous negatives that differ from the positive by only a small detail. We propose Boundary-Aware Curriculum with Local Attention (BACL), a lightweight add-on that turns these borderline cases into a curriculum signal. A Boundary-aware Negative Sampler gradually raises difficulty, while a Contrastive Local Attention loss highlights where the mismatch occurs. The two modules are fully differentiable and work with any off-the-shelf dual encoder. Theory predicts a fast O(1/n) error rate; practice shows up to +32% R@1 over CLIP and new SOTA on four large-scale benchmarks, all without extra labels.
title Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.6; I.2.10; I.2.7
url https://arxiv.org/abs/2511.08399