Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Junxin, Song, Tengfei, Wu, Zhanglin, Li, Pengfei, Liang, Xiaowei, Yang, Hui, Chen, Kun, Xie, Ning, Lu, Yunfei, Zhao, Jing, Sun, Shiliang, Wei, Daimeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918356475445248
author Lu, Junxin
Song, Tengfei
Wu, Zhanglin
Li, Pengfei
Liang, Xiaowei
Yang, Hui
Chen, Kun
Xie, Ning
Lu, Yunfei
Zhao, Jing
Sun, Shiliang
Wei, Daimeng
author_facet Lu, Junxin
Song, Tengfei
Wu, Zhanglin
Li, Pengfei
Liang, Xiaowei
Yang, Hui
Chen, Kun
Xie, Ning
Lu, Yunfei
Zhao, Jing
Sun, Shiliang
Wei, Daimeng
contents Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21956
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
Lu, Junxin
Song, Tengfei
Wu, Zhanglin
Li, Pengfei
Liang, Xiaowei
Yang, Hui
Chen, Kun
Xie, Ning
Lu, Yunfei
Zhao, Jing
Sun, Shiliang
Wei, Daimeng
Computer Vision and Pattern Recognition
Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions.
title Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.21956