DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Zhanglin, Song, Tengfei, Xie, Ning, Zhang, Weidong, Li, Pengfei, Wu, Shuang, Li, Chong, Zhu, Junhao, Yang, Hao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908336057745408
author Wu, Zhanglin
Song, Tengfei
Xie, Ning
Zhang, Weidong
Li, Pengfei
Wu, Shuang
Li, Chong
Zhu, Junhao
Yang, Hao
author_facet Wu, Zhanglin
Song, Tengfei
Xie, Ning
Zhang, Weidong
Li, Pengfei
Wu, Shuang
Li, Chong
Zhu, Junhao
Yang, Hao
contents This paper presents the technical solution proposed by Huawei Translation Service Center (HW-TSC) for the "End-to-End Document Image Machine Translation for Complex Layouts" competition at the 19th International Conference on Document Analysis and Recognition (DIMT25@ICDAR2025). Leveraging state-of-the-art open-source large vision-language model (LVLM), we introduce a training framework that combines multi-task learning with perceptual chain-of-thought to develop a comprehensive end-to-end document translation system. During the inference phase, we apply minimum Bayesian decoding and post-processing strategies to further enhance the system's translation capabilities. Our solution uniquely addresses both OCR-based and OCR-free document image translation tasks within a unified framework. This paper systematically details the training methods, inference strategies, LVLM base models, training data, experimental setups, and results, demonstrating an effective approach to document image machine translation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17315
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model
Wu, Zhanglin
Song, Tengfei
Xie, Ning
Zhang, Weidong
Li, Pengfei
Wu, Shuang
Li, Chong
Zhu, Junhao
Yang, Hao
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper presents the technical solution proposed by Huawei Translation Service Center (HW-TSC) for the "End-to-End Document Image Machine Translation for Complex Layouts" competition at the 19th International Conference on Document Analysis and Recognition (DIMT25@ICDAR2025). Leveraging state-of-the-art open-source large vision-language model (LVLM), we introduce a training framework that combines multi-task learning with perceptual chain-of-thought to develop a comprehensive end-to-end document translation system. During the inference phase, we apply minimum Bayesian decoding and post-processing strategies to further enhance the system's translation capabilities. Our solution uniquely addresses both OCR-based and OCR-free document image translation tasks within a unified framework. This paper systematically details the training methods, inference strategies, LVLM base models, training data, experimental setups, and results, demonstrating an effective approach to document image machine translation.
title DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.17315