Multimodal Machine Translation with Visual Scene Graph Pruning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Chenyu, Sun, Shiliang, Zhao, Jing, Zhang, Nan, Song, Tengfei, Yang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912394754654208
author Lu, Chenyu
Sun, Shiliang
Zhao, Jing
Zhang, Nan
Song, Tengfei
Yang, Hao
author_facet Lu, Chenyu
Sun, Shiliang
Zhao, Jing
Zhang, Nan
Song, Tengfei
Yang, Hao
contents Multimodal machine translation (MMT) seeks to address the challenges posed by linguistic polysemy and ambiguity in translation tasks by incorporating visual information. A key bottleneck in current MMT research is the effective utilization of visual data. Previous approaches have focused on extracting global or region-level image features and using attention or gating mechanisms for multimodal information fusion. However, these methods have not adequately tackled the issue of visual information redundancy in MMT, nor have they proposed effective solutions. In this paper, we introduce a novel approach--multimodal machine translation with visual Scene Graph Pruning (PSG), which leverages language scene graph information to guide the pruning of redundant nodes in visual scene graphs, thereby reducing noise in downstream translation tasks. Through extensive comparative experiments with state-of-the-art methods and ablation studies, we demonstrate the effectiveness of the PSG model. Our results also highlight the promising potential of visual information pruning in advancing the field of MMT.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Machine Translation with Visual Scene Graph Pruning
Lu, Chenyu
Sun, Shiliang
Zhao, Jing
Zhang, Nan
Song, Tengfei
Yang, Hao
Computer Vision and Pattern Recognition
Machine Learning
Multimodal machine translation (MMT) seeks to address the challenges posed by linguistic polysemy and ambiguity in translation tasks by incorporating visual information. A key bottleneck in current MMT research is the effective utilization of visual data. Previous approaches have focused on extracting global or region-level image features and using attention or gating mechanisms for multimodal information fusion. However, these methods have not adequately tackled the issue of visual information redundancy in MMT, nor have they proposed effective solutions. In this paper, we introduce a novel approach--multimodal machine translation with visual Scene Graph Pruning (PSG), which leverages language scene graph information to guide the pruning of redundant nodes in visual scene graphs, thereby reducing noise in downstream translation tasks. Through extensive comparative experiments with state-of-the-art methods and ablation studies, we demonstrate the effectiveness of the PSG model. Our results also highlight the promising potential of visual information pruning in advancing the field of MMT.
title Multimodal Machine Translation with Visual Scene Graph Pruning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.19507