DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jiaxin, Yang, Wentao, Lai, Songxuan, Xie, Zecheng, Jin, Lianwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
von: Liao, Wenhui, et al.
Veröffentlicht: (2024)
von: Liao, Wenhui, et al.
Veröffentlicht: (2024)
Privacy-Preserving Biometric Verification with Handwritten Random Digit String
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
von: Chen, Ketong, et al.
Veröffentlicht: (2025)
von: Chen, Ketong, et al.
Veröffentlicht: (2025)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
LEGO: Self-Supervised Representation Learning for Scene Text Images
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
von: Mo, Ye, et al.
Veröffentlicht: (2025)
von: Mo, Ye, et al.
Veröffentlicht: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
von: Qian, Kun, et al.
Veröffentlicht: (2025)
von: Qian, Kun, et al.
Veröffentlicht: (2025)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
von: Feng, Hao, et al.
Veröffentlicht: (2023)
von: Feng, Hao, et al.
Veröffentlicht: (2023)
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
von: Li, Xin, et al.
Veröffentlicht: (2024)
von: Li, Xin, et al.
Veröffentlicht: (2024)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
von: Saifullah, Saifullah, et al.
Veröffentlicht: (2025)
von: Saifullah, Saifullah, et al.
Veröffentlicht: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
von: Blume, Ansel, et al.
Veröffentlicht: (2025)
von: Blume, Ansel, et al.
Veröffentlicht: (2025)
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
von: Chen, Jian, et al.
Veröffentlicht: (2025)
von: Chen, Jian, et al.
Veröffentlicht: (2025)
Window Token Concatenation for Efficient Visual Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
von: Zhang, Peirong, et al.
Veröffentlicht: (2024)
von: Zhang, Peirong, et al.
Veröffentlicht: (2024)
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
EVLM: An Efficient Vision-Language Model for Visual Understanding
von: Chen, Kaibing, et al.
Veröffentlicht: (2024)
von: Chen, Kaibing, et al.
Veröffentlicht: (2024)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2025)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)
Predicting the Original Appearance of Damaged Historical Documents
von: Yang, Zhenhua, et al.
Veröffentlicht: (2024)
von: Yang, Zhenhua, et al.
Veröffentlicht: (2024)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024)
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024) -
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
von: Liao, Wenhui, et al.
Veröffentlicht: (2024) -
Privacy-Preserving Biometric Verification with Handwritten Random Digit String
von: Zhang, Peirong, et al.
Veröffentlicht: (2025) -
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
von: Chen, Ketong, et al.
Veröffentlicht: (2025) -
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)