Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918356475445248 |
|---|---|
| author | Lu, Junxin Song, Tengfei Wu, Zhanglin Li, Pengfei Liang, Xiaowei Yang, Hui Chen, Kun Xie, Ning Lu, Yunfei Zhao, Jing Sun, Shiliang Wei, Daimeng |
| author_facet | Lu, Junxin Song, Tengfei Wu, Zhanglin Li, Pengfei Liang, Xiaowei Yang, Hui Chen, Kun Xie, Ning Lu, Yunfei Zhao, Jing Sun, Shiliang Wei, Daimeng |
| contents | Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_21956 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation Lu, Junxin Song, Tengfei Wu, Zhanglin Li, Pengfei Liang, Xiaowei Yang, Hui Chen, Kun Xie, Ning Lu, Yunfei Zhao, Jing Sun, Shiliang Wei, Daimeng Computer Vision and Pattern Recognition Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions. |
| title | Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.21956 |