Two Causal Principles for Improving Visual Dialog

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Jiaxin, Niu, Yulei, Huang, Jianqiang, Zhang, Hanwang
Format: Preprint
Published: 2019
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908702359945216
author Qi, Jiaxin
Niu, Yulei
Huang, Jianqiang
Zhang, Hanwang
author_facet Qi, Jiaxin
Niu, Yulei
Huang, Jianqiang
Zhang, Hanwang
contents This paper unravels the design tricks adopted by us, the champion team MReaL-BDAI, for Visual Dialog Challenge 2019: two causal principles for improving Visual Dialog (VisDial). By "improving", we mean that they can promote almost every existing VisDial model to the state-of-the-art performance on the leader-board. Such a major improvement is only due to our careful inspection on the causality behind the model and data, finding that the community has overlooked two causalities in VisDial. Intuitively, Principle 1 suggests: we should remove the direct input of the dialog history to the answer model, otherwise a harmful shortcut bias will be introduced; Principle 2 says: there is an unobserved confounder for history, question, and answer, leading to spurious correlations from training data. In particular, to remove the confounder suggested in Principle 2, we propose several causal intervention algorithms, which make the training fundamentally different from the traditional likelihood estimation. Note that the two principles are model-agnostic, so they are applicable in any VisDial model. The code is available at https://github.com/simpleshinobu/visdial-principles.
format Preprint
id arxiv_https___arxiv_org_abs_1911_10496
institution arXiv
publishDate 2019
record_format arxiv
spellingShingle Two Causal Principles for Improving Visual Dialog
Qi, Jiaxin
Niu, Yulei
Huang, Jianqiang
Zhang, Hanwang
Computer Vision and Pattern Recognition
Computation and Language
This paper unravels the design tricks adopted by us, the champion team MReaL-BDAI, for Visual Dialog Challenge 2019: two causal principles for improving Visual Dialog (VisDial). By "improving", we mean that they can promote almost every existing VisDial model to the state-of-the-art performance on the leader-board. Such a major improvement is only due to our careful inspection on the causality behind the model and data, finding that the community has overlooked two causalities in VisDial. Intuitively, Principle 1 suggests: we should remove the direct input of the dialog history to the answer model, otherwise a harmful shortcut bias will be introduced; Principle 2 says: there is an unobserved confounder for history, question, and answer, leading to spurious correlations from training data. In particular, to remove the confounder suggested in Principle 2, we propose several causal intervention algorithms, which make the training fundamentally different from the traditional likelihood estimation. Note that the two principles are model-agnostic, so they are applicable in any VisDial model. The code is available at https://github.com/simpleshinobu/visdial-principles.
title Two Causal Principles for Improving Visual Dialog
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/1911.10496