Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917197996097536 |
|---|---|
| author | Ye, Hua Ding, Hang Chen, Siyuan Jiang, Yiyang Zhang, Changyuan Zhang, Xuan |
| author_facet | Ye, Hua Ding, Hang Chen, Siyuan Jiang, Yiyang Zhang, Changyuan Zhang, Xuan |
| contents | Most multimodal models treat every negative pair alike, ignoring the ambiguous negatives that differ from the positive by only a small detail. We propose Boundary-Aware Curriculum with Local Attention (BACL), a lightweight add-on that turns these borderline cases into a curriculum signal. A Boundary-aware Negative Sampler gradually raises difficulty, while a Contrastive Local Attention loss highlights where the mismatch occurs. The two modules are fully differentiable and work with any off-the-shelf dual encoder. Theory predicts a fast O(1/n) error rate; practice shows up to +32% R@1 over CLIP and new SOTA on four large-scale benchmarks, all without extra labels. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_08399 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment Ye, Hua Ding, Hang Chen, Siyuan Jiang, Yiyang Zhang, Changyuan Zhang, Xuan Machine Learning Artificial Intelligence Computer Vision and Pattern Recognition I.2.6; I.2.10; I.2.7 Most multimodal models treat every negative pair alike, ignoring the ambiguous negatives that differ from the positive by only a small detail. We propose Boundary-Aware Curriculum with Local Attention (BACL), a lightweight add-on that turns these borderline cases into a curriculum signal. A Boundary-aware Negative Sampler gradually raises difficulty, while a Contrastive Local Attention loss highlights where the mismatch occurs. The two modules are fully differentiable and work with any off-the-shelf dual encoder. Theory predicts a fast O(1/n) error rate; practice shows up to +32% R@1 over CLIP and new SOTA on four large-scale benchmarks, all without extra labels. |
| title | Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment |
| topic | Machine Learning Artificial Intelligence Computer Vision and Pattern Recognition I.2.6; I.2.10; I.2.7 |
| url | https://arxiv.org/abs/2511.08399 |