DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhanglin, Song, Tengfei, Xie, Ning, Zhang, Weidong, Li, Pengfei, Wu, Shuang, Li, Chong, Zhu, Junhao, Yang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HW-TSC's Submission to the CCMT 2024 Machine Translation Tasks
by: Wu, Zhanglin, et al.
Published: (2024)
by: Wu, Zhanglin, et al.
Published: (2024)
Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models
by: Wu, Zhanglin, et al.
Published: (2025)
by: Wu, Zhanglin, et al.
Published: (2025)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
by: Zhang, Yaping, et al.
Published: (2026)
by: Zhang, Yaping, et al.
Published: (2026)
Choose the Final Translation from NMT and LLM hypotheses Using MBR Decoding: HW-TSC's Submission to the WMT24 General MT Shared Task
by: Wu, Zhanglin, et al.
Published: (2024)
by: Wu, Zhanglin, et al.
Published: (2024)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
by: Li, Shaojun, et al.
Published: (2024)
by: Li, Shaojun, et al.
Published: (2024)
An End-to-End Framework for Building Large Language Models for Software Operations
by: He, Jingkai, et al.
Published: (2026)
by: He, Jingkai, et al.
Published: (2026)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
by: Lu, Junxin, et al.
Published: (2026)
by: Lu, Junxin, et al.
Published: (2026)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
by: Guo, Zhao, et al.
Published: (2025)
by: Guo, Zhao, et al.
Published: (2025)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
by: Moslem, Yasmin
Published: (2024)
by: Moslem, Yasmin
Published: (2024)
Team PA-VCG's Solution for Competition on Understanding Chinese College Entrance Exam Papers in ICDAR'25
by: Wu, Wei, et al.
Published: (2025)
by: Wu, Wei, et al.
Published: (2025)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
Taiwan Test Data for ICDAR'25 MapText Competition
by: Lin, Yijun, et al.
Published: (2025)
by: Lin, Yijun, et al.
Published: (2025)
Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation
by: Lan, Zhibin, et al.
Published: (2024)
by: Lan, Zhibin, et al.
Published: (2024)
FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment
by: Li, Xiaohe, et al.
Published: (2025)
by: Li, Xiaohe, et al.
Published: (2025)
Distributionally Robust Control with End-to-End Statistically Guaranteed Metric Learning
by: Wu, Jingyi, et al.
Published: (2025)
by: Wu, Jingyi, et al.
Published: (2025)
Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-ONNs
by: Xie, Yuxin, et al.
Published: (2024)
by: Xie, Yuxin, et al.
Published: (2024)
Dissipative Information Medium Theory (DIMT)
by: Rusishvili, Mikheil
Published: (2026)
by: Rusishvili, Mikheil
Published: (2026)
Dissipative Information Medium Theory (DIMT)
by: Rusishvili, Mikheil
Published: (2026)
by: Rusishvili, Mikheil
Published: (2026)
Dissipative Information Medium Theory (DIMT)
by: Rusishvili, Mikheil
Published: (2026)
by: Rusishvili, Mikheil
Published: (2026)
Towards End-to-End Earthquake Monitoring Using a Multitask Deep Learning Model
by: Zhu, Weiqiang, et al.
Published: (2025)
by: Zhu, Weiqiang, et al.
Published: (2025)
Align-then-Slide: A complete evaluation framework for Ultra-Long Document-Level Machine Translation
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
by: Song, Nan, et al.
Published: (2025)
by: Song, Nan, et al.
Published: (2025)
Exploring the Causality of End-to-End Autonomous Driving
by: Li, Jiankun, et al.
Published: (2024)
by: Li, Jiankun, et al.
Published: (2024)
Vision Transformers for End-to-End Vision-Based Quadrotor Obstacle Avoidance
by: Bhattacharya, Anish, et al.
Published: (2024)
by: Bhattacharya, Anish, et al.
Published: (2024)
Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation
by: Wang, Feifan, et al.
Published: (2025)
by: Wang, Feifan, et al.
Published: (2025)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
by: Lee, Jonghun, et al.
Published: (2026)
by: Lee, Jonghun, et al.
Published: (2026)
Leverage Cross-Attention for End-to-End Open-Vocabulary Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
End-to-End Long Document Summarization using Gradient Caching
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
NORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data Science
by: Zhou, Bing, et al.
Published: (2026)
by: Zhou, Bing, et al.
Published: (2026)
Representation Purification for End-to-End Speech Translation
by: Zhang, Chengwei, et al.
Published: (2024)
by: Zhang, Chengwei, et al.
Published: (2024)
ASCEND: Accurate yet Efficient End-to-End Stochastic Computing Acceleration of Vision Transformer
by: Xie, Tong, et al.
Published: (2024)
by: Xie, Tong, et al.
Published: (2024)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder
by: Dong, Zeyu, et al.
Published: (2026)
by: Dong, Zeyu, et al.
Published: (2026)
Machine Translation Advancements of Low-Resource Indian Languages by Transfer Learning
by: Wei, Bin, et al.
Published: (2024)
by: Wei, Bin, et al.
Published: (2024)
Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving
by: Ge, Junhao, et al.
Published: (2025)
by: Ge, Junhao, et al.
Published: (2025)
Delayed TSC Diagnosis Presenting as End‐Stage Renal Disease With Renal and Hepatic Angiomyolipoma: Case Report and Review
by: Changlin Wei, et al.
Published: (2025)
by: Changlin Wei, et al.
Published: (2025)
GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
by: Liu, Lin, et al.
Published: (2025)
by: Liu, Lin, et al.
Published: (2025)
Similar Items
-
HW-TSC's Submission to the CCMT 2024 Machine Translation Tasks
by: Wu, Zhanglin, et al.
Published: (2024) -
Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models
by: Wu, Zhanglin, et al.
Published: (2025) -
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
by: Zhang, Yaping, et al.
Published: (2026) -
Choose the Final Translation from NMT and LLM hypotheses Using MBR Decoding: HW-TSC's Submission to the WMT24 General MT Shared Task
by: Wu, Zhanglin, et al.
Published: (2024) -
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
by: Li, Shaojun, et al.
Published: (2024)