Exploring OCR-augmented Generation for Bilingual VQA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, JoonHo, Park, Sunho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BEM: Training-Free Background Embedding Memory for False-Positive Suppression in Real-Time Fixed-Background Camera
von: Park, Junwoo, et al.
Veröffentlicht: (2026)
von: Park, Junwoo, et al.
Veröffentlicht: (2026)
Training Unbiased Diffusion Models From Biased Dataset
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
von: Sun, Lin, et al.
Veröffentlicht: (2026)
von: Sun, Lin, et al.
Veröffentlicht: (2026)
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
von: Chen, Song, et al.
Veröffentlicht: (2025)
von: Chen, Song, et al.
Veröffentlicht: (2025)
Integrating Query-aware Segmentation and Cross-Attention for Robust VQA
von: Choi, Wonjun, et al.
Veröffentlicht: (2024)
von: Choi, Wonjun, et al.
Veröffentlicht: (2024)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
von: Liu, Yibo, et al.
Veröffentlicht: (2024)
von: Liu, Yibo, et al.
Veröffentlicht: (2024)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
von: Wen, Shimin, et al.
Veröffentlicht: (2026)
von: Wen, Shimin, et al.
Veröffentlicht: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
Enhancing Document VQA Models via Retrieval-Augmented Generation
von: López, Eric, et al.
Veröffentlicht: (2025)
von: López, Eric, et al.
Veröffentlicht: (2025)
SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models
von: Park, Joon Hyun, et al.
Veröffentlicht: (2025)
von: Park, Joon Hyun, et al.
Veröffentlicht: (2025)
On the Role of Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment
von: Jin, Jianing, et al.
Veröffentlicht: (2025)
von: Jin, Jianing, et al.
Veröffentlicht: (2025)
Agentar-Fin-OCR
von: Qian, Siyi, et al.
Veröffentlicht: (2026)
von: Qian, Siyi, et al.
Veröffentlicht: (2026)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2024)
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
VQA Training Sets are Self-play Environments for Generating Few-shot Pools
von: Misiunas, Tautvydas, et al.
Veröffentlicht: (2024)
von: Misiunas, Tautvydas, et al.
Veröffentlicht: (2024)
EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
von: Chen, Baoliang, et al.
Veröffentlicht: (2026)
von: Chen, Baoliang, et al.
Veröffentlicht: (2026)
TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2024)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2024)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
ABot-OCR Technical Report
von: Jiang, Kaitao, et al.
Veröffentlicht: (2026)
von: Jiang, Kaitao, et al.
Veröffentlicht: (2026)
Post-Training Quantization via Residual Truncation and Zero Suppression for Diffusion Models
von: Kim, Donghoon, et al.
Veröffentlicht: (2025)
von: Kim, Donghoon, et al.
Veröffentlicht: (2025)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
von: Kashid, Harshvivek, et al.
Veröffentlicht: (2024)
von: Kashid, Harshvivek, et al.
Veröffentlicht: (2024)
Elevating Flow-Guided Video Inpainting with Reference Generation
von: Cho, Suhwan, et al.
Veröffentlicht: (2024)
von: Cho, Suhwan, et al.
Veröffentlicht: (2024)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
von: Li, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2023)
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
von: Vesalainen, Ari, et al.
Veröffentlicht: (2026)
von: Vesalainen, Ari, et al.
Veröffentlicht: (2026)
Knowledge Condensation and Reasoning for Knowledge-based VQA
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
Measuring Faithful and Plausible Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2023)
von: Reich, Daniel, et al.
Veröffentlicht: (2023)
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
von: Gong, Weile, et al.
Veröffentlicht: (2026)
von: Gong, Weile, et al.
Veröffentlicht: (2026)
Exploring Temporally-Aware Features for Point Tracking
von: Kim, Inès Hyeonsu, et al.
Veröffentlicht: (2025)
von: Kim, Inès Hyeonsu, et al.
Veröffentlicht: (2025)
DODO: Discrete OCR Diffusion Models
von: Man, Sean, et al.
Veröffentlicht: (2026)
von: Man, Sean, et al.
Veröffentlicht: (2026)
An Empirical Study of Scaling Law for OCR
von: Rang, Miao, et al.
Veröffentlicht: (2023)
von: Rang, Miao, et al.
Veröffentlicht: (2023)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation
von: Chen, Bowen, et al.
Veröffentlicht: (2025)
von: Chen, Bowen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BEM: Training-Free Background Embedding Memory for False-Positive Suppression in Real-Time Fixed-Background Camera
von: Park, Junwoo, et al.
Veröffentlicht: (2026) -
Training Unbiased Diffusion Models From Biased Dataset
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024) -
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024) -
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
von: Sun, Lin, et al.
Veröffentlicht: (2026) -
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
von: Park, Jaeyoo, et al.
Veröffentlicht: (2024)