MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Gengluo, Zhang, Chengquan, Liang, Yupu, Shen, Huawen, Zhang, Yaping, Lyu, Pengyuan, Wang, Weinong, Wan, Xingyu, Zeng, Gangyan, Hu, Han, Ma, Can, Zhou, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
von: Li, Gengluo, et al.
Veröffentlicht: (2025)
von: Li, Gengluo, et al.
Veröffentlicht: (2025)
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
von: Zhang, Yan, et al.
Veröffentlicht: (2024)
von: Zhang, Yan, et al.
Veröffentlicht: (2024)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
Towards Unified Multi-granularity Text Detection with Interactive Attention
von: Wan, Xingyu, et al.
Veröffentlicht: (2024)
von: Wan, Xingyu, et al.
Veröffentlicht: (2024)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024)
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024)
Recognition-Synergistic Scene Text Editing
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting
von: Wu, Jingjing, et al.
Veröffentlicht: (2024)
von: Wu, Jingjing, et al.
Veröffentlicht: (2024)
Falcon-UI: Understanding GUI Before Following User Instructions
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
von: Lyu, Jiahao, et al.
Veröffentlicht: (2026)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2026)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
von: Zhang, Yaping, et al.
Veröffentlicht: (2025)
von: Zhang, Yaping, et al.
Veröffentlicht: (2025)
Towards Training-Free Scene Text Editing
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
HunyuanOCR Technical Report
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Solving the Fleet Size and Mix Vehicle Routing Problem
von: Wan, Pengfu, et al.
Veröffentlicht: (2025)
von: Wan, Pengfu, et al.
Veröffentlicht: (2025)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
von: Zhang, Yaping, et al.
Veröffentlicht: (2026)
von: Zhang, Yaping, et al.
Veröffentlicht: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
von: Lee, Yunseung, et al.
Veröffentlicht: (2026)
von: Lee, Yunseung, et al.
Veröffentlicht: (2026)
PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain
von: Jin, Song, et al.
Veröffentlicht: (2026)
von: Jin, Song, et al.
Veröffentlicht: (2026)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
von: Shi, Enyi, et al.
Veröffentlicht: (2026)
von: Shi, Enyi, et al.
Veröffentlicht: (2026)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark for Text Categorization
von: Costache, Adriana-Valentina, et al.
Veröffentlicht: (2026)
von: Costache, Adriana-Valentina, et al.
Veröffentlicht: (2026)
AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
von: Zha, Jirong, et al.
Veröffentlicht: (2025)
von: Zha, Jirong, et al.
Veröffentlicht: (2025)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
von: Chen, Ruihan, et al.
Veröffentlicht: (2025)
von: Chen, Ruihan, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
TransBench: Benchmarking Machine Translation for Industrial-Scale Applications
von: Li, Haijun, et al.
Veröffentlicht: (2025)
von: Li, Haijun, et al.
Veröffentlicht: (2025)
VideoMarkBench: Benchmarking Robustness of Video Watermarking
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
von: Xun, Shuhang, et al.
Veröffentlicht: (2025)
von: Xun, Shuhang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026) -
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
von: Li, Gengluo, et al.
Veröffentlicht: (2025) -
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
von: Zhang, Yan, et al.
Veröffentlicht: (2024) -
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
von: Shen, Huawen, et al.
Veröffentlicht: (2024) -
Towards Unified Multi-granularity Text Detection with Interactive Attention
von: Wan, Xingyu, et al.
Veröffentlicht: (2024)