TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yuliang, Yang, Biao, Liu, Qiang, Li, Zhang, Ma, Zhiyin, Zhang, Shuo, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
von: Li, Zhang, et al.
Veröffentlicht: (2023)
von: Li, Zhang, et al.
Veröffentlicht: (2023)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
Exploring the Capabilities of Large Multimodal Models on Dense Text
von: Zhang, Shuo, et al.
Veröffentlicht: (2024)
von: Zhang, Shuo, et al.
Veröffentlicht: (2024)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
Multimodal OCR: Parse Anything from Documents
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
von: Fu, Ling, et al.
Veröffentlicht: (2024)
von: Fu, Ling, et al.
Veröffentlicht: (2024)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
von: Wasfy, Ahmed, et al.
Veröffentlicht: (2025)
von: Wasfy, Ahmed, et al.
Veröffentlicht: (2025)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
von: Haq, Ijazul, et al.
Veröffentlicht: (2025)
von: Haq, Ijazul, et al.
Veröffentlicht: (2025)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
von: Feng, Hao, et al.
Veröffentlicht: (2023)
von: Feng, Hao, et al.
Veröffentlicht: (2023)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
von: Tang, Shuo, et al.
Veröffentlicht: (2025)
von: Tang, Shuo, et al.
Veröffentlicht: (2025)
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)
TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2024)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2024)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2023)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2023)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
von: Yin, Liang, et al.
Veröffentlicht: (2025)
von: Yin, Liang, et al.
Veröffentlicht: (2025)
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
von: Yan, Hao, et al.
Veröffentlicht: (2025)
von: Yan, Hao, et al.
Veröffentlicht: (2025)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
von: Zhang, Yulong
Veröffentlicht: (2025)
von: Zhang, Yulong
Veröffentlicht: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
WAS: Dataset and Methods for Artistic Text Segmentation
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
von: Tien, Dong Nguyen, et al.
Veröffentlicht: (2025)
von: Tien, Dong Nguyen, et al.
Veröffentlicht: (2025)
Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
von: Li, Yinghui, et al.
Veröffentlicht: (2026)
von: Li, Yinghui, et al.
Veröffentlicht: (2026)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
von: Chern, Ethan, et al.
Veröffentlicht: (2024)
von: Chern, Ethan, et al.
Veröffentlicht: (2024)
HunyuanOCR Technical Report
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
von: Li, Zhang, et al.
Veröffentlicht: (2023) -
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
von: Li, Zhang, et al.
Veröffentlicht: (2025) -
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
von: Li, Zhang, et al.
Veröffentlicht: (2025) -
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025) -
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)