LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Jingxuan, Jia, Caijun, Chen, Qi, Cai, Yujun, Sun, Linzhuang, Zhang, Xiangxiang, Wu, Gaowei, Yu, Bihui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912500976451584
author Wei, Jingxuan
Jia, Caijun
Chen, Qi
Cai, Yujun
Sun, Linzhuang
Zhang, Xiangxiang
Wu, Gaowei
Yu, Bihui
author_facet Wei, Jingxuan
Jia, Caijun
Chen, Qi
Cai, Yujun
Sun, Linzhuang
Zhang, Xiangxiang
Wu, Gaowei
Yu, Bihui
contents Multimodal Machine Translation (MMT) enhances translation quality by incorporating visual context, helping to resolve textual ambiguities. While existing MMT methods perform well in bilingual settings, extending them to multilingual translation remains challenging due to cross-lingual interference and ineffective parameter-sharing strategies. To address this, we propose LLaVA-NeuMT, a novel multimodal multilingual translation framework that explicitly models language-specific and language-agnostic representations to mitigate multilingual interference. Our approach consists of a layer selection mechanism that identifies the most informative layers for different language pairs and a neuron-level adaptation strategy that dynamically selects language-specific and agnostic neurons to improve translation quality while reducing redundancy. We conduct extensive experiments on the M3-Multi30K and M3-AmbigCaps datasets, demonstrating that LLaVA-NeuMT, while fine-tuning only 40\% of the model parameters, surpasses full fine-tuning approaches and ultimately achieves SOTA results on both datasets. Our analysis further provides insights into the importance of selected layers and neurons in multimodal multilingual adaptation, offering an efficient and scalable solution to cross-lingual adaptation in multimodal translation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
Wei, Jingxuan
Jia, Caijun
Chen, Qi
Cai, Yujun
Sun, Linzhuang
Zhang, Xiangxiang
Wu, Gaowei
Yu, Bihui
Computation and Language
Multimedia
Multimodal Machine Translation (MMT) enhances translation quality by incorporating visual context, helping to resolve textual ambiguities. While existing MMT methods perform well in bilingual settings, extending them to multilingual translation remains challenging due to cross-lingual interference and ineffective parameter-sharing strategies. To address this, we propose LLaVA-NeuMT, a novel multimodal multilingual translation framework that explicitly models language-specific and language-agnostic representations to mitigate multilingual interference. Our approach consists of a layer selection mechanism that identifies the most informative layers for different language pairs and a neuron-level adaptation strategy that dynamically selects language-specific and agnostic neurons to improve translation quality while reducing redundancy. We conduct extensive experiments on the M3-Multi30K and M3-AmbigCaps datasets, demonstrating that LLaVA-NeuMT, while fine-tuning only 40\% of the model parameters, surpasses full fine-tuning approaches and ultimately achieves SOTA results on both datasets. Our analysis further provides insights into the importance of selected layers and neurons in multimodal multilingual adaptation, offering an efficient and scalable solution to cross-lingual adaptation in multimodal translation.
title LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2507.18940