TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akl, Ahmed, Khamis, Abdelwahed, Wang, Zhe, Cheraghian, Ali, Khalifa, Sara, Wang, Kewen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918402314993664
author Akl, Ahmed
Khamis, Abdelwahed
Wang, Zhe
Cheraghian, Ali
Khalifa, Sara
Wang, Kewen
author_facet Akl, Ahmed
Khamis, Abdelwahed
Wang, Zhe
Cheraghian, Ali
Khalifa, Sara
Wang, Kewen
contents Visual Question Answering (VQA) systems are notoriously brittle under distribution shifts and data scarcity. While previous solutions-such as ensemble methods and data augmentation-can improve performance in isolation, they fail to generalise well across in-distribution (IID), out-of-distribution (OOD), and low-data settings simultaneously. We argue that this limitation stems from the suboptimal training strategies employed. Specifically, treating all training samples uniformly-without accounting for question difficulty or semantic structure-leaves the models vulnerable to dataset biases. Thus, they struggle to generalise beyond the training distribution. To address this issue, we introduce Task-Progressive Curriculum Learning (TPCL)-a simple, model-agnostic framework that progressively trains VQA models using a curriculum built by jointly considering question type and difficulty. Specifically, TPCL first groups questions based on their semantic type (e.g., yes/no, counting) and then orders them using a novel Optimal Transport-based difficulty measure. Without relying on data augmentation or explicit debiasing, TPCL improves generalisation across IID, OOD, and low-data regimes and achieves state-of-the-art performance on VQA-CP v2, VQA-CP v1, and VQA v2. It outperforms the most competitive robust VQA baselines by over 5% and 7% on VQA-CP v2 and v1, respectively, and boosts backbone performance by up to 28.5%.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
Akl, Ahmed
Khamis, Abdelwahed
Wang, Zhe
Cheraghian, Ali
Khalifa, Sara
Wang, Kewen
Computer Vision and Pattern Recognition
Machine Learning
Visual Question Answering (VQA) systems are notoriously brittle under distribution shifts and data scarcity. While previous solutions-such as ensemble methods and data augmentation-can improve performance in isolation, they fail to generalise well across in-distribution (IID), out-of-distribution (OOD), and low-data settings simultaneously. We argue that this limitation stems from the suboptimal training strategies employed. Specifically, treating all training samples uniformly-without accounting for question difficulty or semantic structure-leaves the models vulnerable to dataset biases. Thus, they struggle to generalise beyond the training distribution. To address this issue, we introduce Task-Progressive Curriculum Learning (TPCL)-a simple, model-agnostic framework that progressively trains VQA models using a curriculum built by jointly considering question type and difficulty. Specifically, TPCL first groups questions based on their semantic type (e.g., yes/no, counting) and then orders them using a novel Optimal Transport-based difficulty measure. Without relying on data augmentation or explicit debiasing, TPCL improves generalisation across IID, OOD, and low-data regimes and achieves state-of-the-art performance on VQA-CP v2, VQA-CP v1, and VQA v2. It outperforms the most competitive robust VQA baselines by over 5% and 7% on VQA-CP v2 and v1, respectively, and boosts backbone performance by up to 28.5%.
title TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.17292