Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913045950758912 |
|---|---|
| author | Li, Jiangnan Huang, Linqing Yan, Xiaowen Gan, Min Lu, Wenpeng Fan, Jinfu |
| author_facet | Li, Jiangnan Huang, Linqing Yan, Xiaowen Gan, Min Lu, Wenpeng Fan, Jinfu |
| contents | Zero-shot learning (ZSL) aims to recognize unseen classes without visual instances. However, existing methods usually assume clean labels, overlooking real-world label noise and ambiguity, which degrades performance. To bridge this gap, we propose the Dynamic Visual-semantic Alignment (DVSA), a robust ZSL framework for learning from ambiguous labels. DVSA uses a bidirectional visual-semantic alignment module with attention to mutually calibrate visual features and attribute prototypes, and a contrastive optimization grounded in Mutual Information (MI) at the attribute level to strengthen discriminative, semantically consistent attributes. In addition, a dynamic label disambiguation mechanism iteratively corrects noisy supervision while preserving semantic consistency, narrowing the instance-label gap, and improving generalization. Extensive experiments on standard benchmarks verify that DVSA achieves stronger performance under ambiguous supervision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_17710 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels Li, Jiangnan Huang, Linqing Yan, Xiaowen Gan, Min Lu, Wenpeng Fan, Jinfu Computer Vision and Pattern Recognition Zero-shot learning (ZSL) aims to recognize unseen classes without visual instances. However, existing methods usually assume clean labels, overlooking real-world label noise and ambiguity, which degrades performance. To bridge this gap, we propose the Dynamic Visual-semantic Alignment (DVSA), a robust ZSL framework for learning from ambiguous labels. DVSA uses a bidirectional visual-semantic alignment module with attention to mutually calibrate visual features and attribute prototypes, and a contrastive optimization grounded in Mutual Information (MI) at the attribute level to strengthen discriminative, semantically consistent attributes. In addition, a dynamic label disambiguation mechanism iteratively corrects noisy supervision while preserving semantic consistency, narrowing the instance-label gap, and improving generalization. Extensive experiments on standard benchmarks verify that DVSA achieves stronger performance under ambiguous supervision. |
| title | Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.17710 |