Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Yupu, Zhang, Yaping, Zhang, Zhiyang, Chen, Zhiyuan, Zhao, Yang, Xiang, Lu, Zong, Chengqing, Zhou, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
von: Zhang, Yaping, et al.
Veröffentlicht: (2026)
von: Zhang, Yaping, et al.
Veröffentlicht: (2026)
From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignment
von: Ye, Jing, et al.
Veröffentlicht: (2025)
von: Ye, Jing, et al.
Veröffentlicht: (2025)
A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
von: Xiang, Lu, et al.
Veröffentlicht: (2025)
von: Xiang, Lu, et al.
Veröffentlicht: (2025)
EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal World
von: Ye, Jing, et al.
Veröffentlicht: (2026)
von: Ye, Jing, et al.
Veröffentlicht: (2026)
SweetieChat: A Strategy-Enhanced Role-playing Framework for Diverse Scenarios Handling Emotional Support Agent
von: Ye, Jing, et al.
Veröffentlicht: (2024)
von: Ye, Jing, et al.
Veröffentlicht: (2024)
HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
von: Zhang, Yaping, et al.
Veröffentlicht: (2025)
von: Zhang, Yaping, et al.
Veröffentlicht: (2025)
Self-Modifying State Modeling for Simultaneous Machine Translation
von: Yu, Donglei, et al.
Veröffentlicht: (2024)
von: Yu, Donglei, et al.
Veröffentlicht: (2024)
PromptDLA: A Domain-aware Prompt Document Layout Analysis Framework with Descriptive Knowledge as a Cue
von: Zhang, Zirui, et al.
Veröffentlicht: (2026)
von: Zhang, Zirui, et al.
Veröffentlicht: (2026)
F-MALLOC: Feed-forward Memory Allocation for Continual Learning in Neural Machine Translation
von: Wu, Junhong, et al.
Veröffentlicht: (2024)
von: Wu, Junhong, et al.
Veröffentlicht: (2024)
SimulPL: Aligning Human Preferences in Simultaneous Machine Translation
von: Yu, Donglei, et al.
Veröffentlicht: (2025)
von: Yu, Donglei, et al.
Veröffentlicht: (2025)
MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment
von: Li, Chong, et al.
Veröffentlicht: (2023)
von: Li, Chong, et al.
Veröffentlicht: (2023)
Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation
von: Wu, Junhong, et al.
Veröffentlicht: (2024)
von: Wu, Junhong, et al.
Veröffentlicht: (2024)
TokAlign: Efficient Vocabulary Adaptation via Token Alignment
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
von: Ye, Chunyu, et al.
Veröffentlicht: (2025)
von: Ye, Chunyu, et al.
Veröffentlicht: (2025)
Language Imbalance Driven Rewarding for Multilingual Self-improving
von: Yang, Wen, et al.
Veröffentlicht: (2024)
von: Yang, Wen, et al.
Veröffentlicht: (2024)
X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions
von: Li, Chong, et al.
Veröffentlicht: (2024)
von: Li, Chong, et al.
Veröffentlicht: (2024)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
Listening to the Echo: User-Reaction Aware Policy Optimization via Scalar-Verbal Hybrid Reinforcement Learning
von: Ye, Jing, et al.
Veröffentlicht: (2026)
von: Ye, Jing, et al.
Veröffentlicht: (2026)
Dual-branch Prompting for Multimodal Machine Translation
von: Wang, Jie, et al.
Veröffentlicht: (2025)
von: Wang, Jie, et al.
Veröffentlicht: (2025)
Multimodal OCR: Parse Anything from Documents
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models
von: Zhang, Yunhao, et al.
Veröffentlicht: (2024)
von: Zhang, Yunhao, et al.
Veröffentlicht: (2024)
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR
von: Yuan, Ziye, et al.
Veröffentlicht: (2026)
von: Yuan, Ziye, et al.
Veröffentlicht: (2026)
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
von: Yang, Yuwei, et al.
Veröffentlicht: (2025)
von: Yang, Yuwei, et al.
Veröffentlicht: (2025)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
von: Xiao, Han, et al.
Veröffentlicht: (2025)
von: Xiao, Han, et al.
Veröffentlicht: (2025)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
von: Huang, Han, et al.
Veröffentlicht: (2024)
von: Huang, Han, et al.
Veröffentlicht: (2024)
Improve MLLM Benchmark Efficiency through Interview
von: Wen, Farong, et al.
Veröffentlicht: (2025)
von: Wen, Farong, et al.
Veröffentlicht: (2025)
Improving LLM-based Document-level Machine Translation with Multi-Knowledge Fusion
von: Liu, Bin, et al.
Veröffentlicht: (2025)
von: Liu, Bin, et al.
Veröffentlicht: (2025)
Navigating Brain Language Representations: A Comparative Analysis of Neural Language Models and Psychologically Plausible Models
von: Zhang, Yunhao, et al.
Veröffentlicht: (2024)
von: Zhang, Yunhao, et al.
Veröffentlicht: (2024)
TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment
von: Li, Chong, et al.
Veröffentlicht: (2026)
von: Li, Chong, et al.
Veröffentlicht: (2026)
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
von: Wang, Chengye, et al.
Veröffentlicht: (2026)
von: Wang, Chengye, et al.
Veröffentlicht: (2026)
GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
von: Lan, Xiang, et al.
Veröffentlicht: (2025)
von: Lan, Xiang, et al.
Veröffentlicht: (2025)
Improving MLLM Historical Record Extraction with Test-Time Image
von: Archibald, Taylor, et al.
Veröffentlicht: (2025)
von: Archibald, Taylor, et al.
Veröffentlicht: (2025)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
MapGuide: A Simple yet Effective Method to Reconstruct Continuous Language from Brain Activities
von: Zhao, Xinpei, et al.
Veröffentlicht: (2024)
von: Zhao, Xinpei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
von: Liang, Yupu, et al.
Veröffentlicht: (2025) -
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
von: Zhang, Yaping, et al.
Veröffentlicht: (2026) -
From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignment
von: Ye, Jing, et al.
Veröffentlicht: (2025) -
A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
von: Xiang, Lu, et al.
Veröffentlicht: (2025) -
EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal World
von: Ye, Jing, et al.
Veröffentlicht: (2026)