Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Ying, Bai, Ge, Lu, Chenji, Li, Shilong, Zhang, Zhang, Liu, Ruifang, Guo, Wenbin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2410.10184
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910648882954240
author Liu, Ying
Bai, Ge
Lu, Chenji
Li, Shilong
Zhang, Zhang
Liu, Ruifang
Guo, Wenbin
author_facet Liu, Ying
Bai, Ge
Lu, Chenji
Li, Shilong
Zhang, Zhang
Liu, Ruifang
Guo, Wenbin
contents Despite the remarkable advancements in Visual Question Answering (VQA), the challenge of mitigating the language bias introduced by textual information remains unresolved. Previous approaches capture language bias from a coarse-grained perspective. However, the finer-grained information within a sentence, such as context and keywords, can result in different biases. Due to the ignorance of fine-grained information, most existing methods fail to sufficiently capture language bias. In this paper, we propose a novel causal intervention training scheme named CIBi to eliminate language bias from a finer-grained perspective. Specifically, we divide the language bias into context bias and keyword bias. We employ causal intervention and contrastive learning to eliminate context bias and improve the multi-modal representation. Additionally, we design a new question-only branch based on counterfactual generation to distill and eliminate keyword bias. Experimental results illustrate that CIBi is applicable to various VQA models, yielding competitive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
Liu, Ying
Bai, Ge
Lu, Chenji
Li, Shilong
Zhang, Zhang
Liu, Ruifang
Guo, Wenbin
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite the remarkable advancements in Visual Question Answering (VQA), the challenge of mitigating the language bias introduced by textual information remains unresolved. Previous approaches capture language bias from a coarse-grained perspective. However, the finer-grained information within a sentence, such as context and keywords, can result in different biases. Due to the ignorance of fine-grained information, most existing methods fail to sufficiently capture language bias. In this paper, we propose a novel causal intervention training scheme named CIBi to eliminate language bias from a finer-grained perspective. Specifically, we divide the language bias into context bias and keyword bias. We employ causal intervention and contrastive learning to eliminate context bias and improve the multi-modal representation. Additionally, we design a new question-only branch based on counterfactual generation to distill and eliminate keyword bias. Experimental results illustrate that CIBi is applicable to various VQA models, yielding competitive performance.
title Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.10184