Video-guided Machine Translation with Global Video Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jian, Lv, JinZe, Long, Zi, Fu, XiangHua
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913015326048256
author Chen, Jian
Lv, JinZe
Long, Zi
Fu, XiangHua
author_facet Chen, Jian
Lv, JinZe
Long, Zi
Fu, XiangHua
contents Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capture global narrative context across multiple segments in long videos. To overcome this limitation, we propose a globally video-guided multimodal translation framework that leverages a pretrained semantic encoder and vector database-based subtitle retrieval to construct a context set of video segments closely related to the target subtitle semantics. An attention mechanism is employed to focus on highly relevant visual content, while preserving the remaining video features to retain broader contextual information. Furthermore, we design a region-aware cross-modal attention mechanism to enhance semantic alignment during translation. Experiments on a large-scale documentary translation dataset demonstrate that our method significantly outperforms baseline models, highlighting its effectiveness in long-video scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Video-guided Machine Translation with Global Video Context
Chen, Jian
Lv, JinZe
Long, Zi
Fu, XiangHua
Computer Vision and Pattern Recognition
Computation and Language
Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capture global narrative context across multiple segments in long videos. To overcome this limitation, we propose a globally video-guided multimodal translation framework that leverages a pretrained semantic encoder and vector database-based subtitle retrieval to construct a context set of video segments closely related to the target subtitle semantics. An attention mechanism is employed to focus on highly relevant visual content, while preserving the remaining video features to retain broader contextual information. Furthermore, we design a region-aware cross-modal attention mechanism to enhance semantic alignment during translation. Experiments on a large-scale documentary translation dataset demonstrate that our method significantly outperforms baseline models, highlighting its effectiveness in long-video scenarios.
title Video-guided Machine Translation with Global Video Context
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.06789