Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
Fuente:
arXiv
Saved in:
| Main Authors: | Vempati, Shashank, Anand, Nishit, Talebailkar, Gaurav, Garai, Arpan, Arora, Chetan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024)
by: Li, Bozhou, et al.
Published: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
by: Tang, Raphael, et al.
Published: (2024)
by: Tang, Raphael, et al.
Published: (2024)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
by: Guan, Shuhao, et al.
Published: (2025)
by: Guan, Shuhao, et al.
Published: (2025)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
by: Kargaran, Amir Hossein, et al.
Published: (2026)
by: Kargaran, Amir Hossein, et al.
Published: (2026)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
by: Gagnier, Henry, et al.
Published: (2026)
by: Gagnier, Henry, et al.
Published: (2026)
Towards Deployable OCR models for Indic languages
by: Mathew, Minesh, et al.
Published: (2022)
by: Mathew, Minesh, et al.
Published: (2022)
Improving OCR for Historical Texts of Multiple Languages
by: Westerdijk, Hylke, et al.
Published: (2025)
by: Westerdijk, Hylke, et al.
Published: (2025)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024)
by: Kashid, Harshvivek, et al.
Published: (2024)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
by: Liu, Yuliang, et al.
Published: (2023)
by: Liu, Yuliang, et al.
Published: (2023)
Seeing Straight: Document Orientation Detection for Efficient OCR
by: Goswami, Suranjan, et al.
Published: (2025)
by: Goswami, Suranjan, et al.
Published: (2025)
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
by: Raj, Ankita, et al.
Published: (2025)
by: Raj, Ankita, et al.
Published: (2025)
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
by: Phukan, Arpan, et al.
Published: (2024)
by: Phukan, Arpan, et al.
Published: (2024)
ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
by: Abdallah, Abdelrahman, et al.
Published: (2024)
by: Abdallah, Abdelrahman, et al.
Published: (2024)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
by: Hennara, Khalil, et al.
Published: (2025)
by: Hennara, Khalil, et al.
Published: (2025)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
by: Jung, Jaeyoon, et al.
Published: (2026)
by: Jung, Jaeyoon, et al.
Published: (2026)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
by: Liang, Yunhao, et al.
Published: (2026)
by: Liang, Yunhao, et al.
Published: (2026)
Advances and Limitations in Open Source Arabic-Script OCR: A Case Study
by: Kiessling, Benjamin, et al.
Published: (2024)
by: Kiessling, Benjamin, et al.
Published: (2024)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
by: Wang, Zhengren, et al.
Published: (2026)
by: Wang, Zhengren, et al.
Published: (2026)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
by: Cardoso, Gabriel Pimenta de Freitas, et al.
Published: (2026)
by: Cardoso, Gabriel Pimenta de Freitas, et al.
Published: (2026)
Reliable Active Learning from Unreliable Labels via Neural Collapse Geometry
by: Goel, Atharv, et al.
Published: (2025)
by: Goel, Atharv, et al.
Published: (2025)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
by: Poggi, Nicolas, et al.
Published: (2025)
by: Poggi, Nicolas, et al.
Published: (2025)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024)
by: Do, Thao, et al.
Published: (2024)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
by: Tang, Zihan, et al.
Published: (2026)
by: Tang, Zihan, et al.
Published: (2026)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion
by: Jindal, Akshit, et al.
Published: (2026)
by: Jindal, Akshit, et al.
Published: (2026)
A Grounded Typology of Word Classes
by: Haley, Coleman, et al.
Published: (2024)
by: Haley, Coleman, et al.
Published: (2024)
Confidence-Aware Document OCR Error Detection
by: Hemmer, Arthur, et al.
Published: (2024)
by: Hemmer, Arthur, et al.
Published: (2024)
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
by: Le, Binh M., et al.
Published: (2025)
by: Le, Binh M., et al.
Published: (2025)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
by: Al-Homoud, Haneen, et al.
Published: (2025)
by: Al-Homoud, Haneen, et al.
Published: (2025)
PubMed-OCR: PMC Open Access OCR Annotations
by: Heidenreich, Hunter, et al.
Published: (2026)
by: Heidenreich, Hunter, et al.
Published: (2026)
Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs
by: Dey, Abhishek, et al.
Published: (2025)
by: Dey, Abhishek, et al.
Published: (2025)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
by: Madan, Chetan, et al.
Published: (2024)
by: Madan, Chetan, et al.
Published: (2024)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
by: Garg, Roopal, et al.
Published: (2024)
by: Garg, Roopal, et al.
Published: (2024)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
by: Goel, Arushi, et al.
Published: (2026)
by: Goel, Arushi, et al.
Published: (2026)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Why Instruction-Based Unlearning Fails in Diffusion Models?
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
Similar Items
-
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024) -
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025) -
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
by: Tang, Raphael, et al.
Published: (2024) -
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
by: Guan, Shuhao, et al.
Published: (2025) -
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
by: Kargaran, Amir Hossein, et al.
Published: (2026)