Classifier-guided Gradient Modulation for Enhanced Multimodal Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Zirun, Jin, Tao, Chen, Jingyuan, Zhao, Zhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916466889064448
author Guo, Zirun
Jin, Tao
Chen, Jingyuan
Zhao, Zhou
author_facet Guo, Zirun
Jin, Tao
Chen, Jingyuan
Zhao, Zhou
contents Multimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other modalities. Existing methods to balance the training process always have some limitations on the loss functions, optimizers and the number of modalities and only consider modulating the magnitude of the gradients while ignoring the directions of the gradients. To solve these problems, in this paper, we present a novel method to balance multimodal learning with Classifier-Guided Gradient Modulation (CGGM), considering both the magnitude and directions of the gradients. We conduct extensive experiments on four multimodal datasets: UPMC-Food 101, CMU-MOSI, IEMOCAP and BraTS 2021, covering classification, regression and segmentation tasks. The results show that CGGM outperforms all the baselines and other state-of-the-art methods consistently, demonstrating its effectiveness and versatility. Our code is available at https://github.com/zrguo/CGGM.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01409
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
Guo, Zirun
Jin, Tao
Chen, Jingyuan
Zhao, Zhou
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Multimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other modalities. Existing methods to balance the training process always have some limitations on the loss functions, optimizers and the number of modalities and only consider modulating the magnitude of the gradients while ignoring the directions of the gradients. To solve these problems, in this paper, we present a novel method to balance multimodal learning with Classifier-Guided Gradient Modulation (CGGM), considering both the magnitude and directions of the gradients. We conduct extensive experiments on four multimodal datasets: UPMC-Food 101, CMU-MOSI, IEMOCAP and BraTS 2021, covering classification, regression and segmentation tasks. The results show that CGGM outperforms all the baselines and other state-of-the-art methods consistently, demonstrating its effectiveness and versatility. Our code is available at https://github.com/zrguo/CGGM.
title Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.01409