NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Ulla, Aman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
von: Horn, Pius, et al.
Veröffentlicht: (2025)
von: Horn, Pius, et al.
Veröffentlicht: (2025)
From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline
von: Shoilee, Sarah Binta Alam, et al.
Veröffentlicht: (2026)
von: Shoilee, Sarah Binta Alam, et al.
Veröffentlicht: (2026)
Beyond String Matching: Semantic Evaluation of PDF Table Extraction
von: Horn, Pius, et al.
Veröffentlicht: (2026)
von: Horn, Pius, et al.
Veröffentlicht: (2026)
AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
von: Patel, Piyushkumar
Veröffentlicht: (2025)
von: Patel, Piyushkumar
Veröffentlicht: (2025)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items
von: Lin, Jianghao, et al.
Veröffentlicht: (2025)
von: Lin, Jianghao, et al.
Veröffentlicht: (2025)
Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
TalentMine: LLM-Based Extraction and Question-Answering from Multimodal Talent Tables
von: Mannam, Varun, et al.
Veröffentlicht: (2025)
von: Mannam, Varun, et al.
Veröffentlicht: (2025)
CFIR: Fast and Effective Long-Text To Image Retrieval for Large Corpora
von: Long, Zijun, et al.
Veröffentlicht: (2024)
von: Long, Zijun, et al.
Veröffentlicht: (2024)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
von: Lim, Ho Hung, et al.
Veröffentlicht: (2026)
von: Lim, Ho Hung, et al.
Veröffentlicht: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
Good Scores, Bad Data: A Metric for Multimodal Coherence
von: Srinivasan, Vasundra
Veröffentlicht: (2026)
von: Srinivasan, Vasundra
Veröffentlicht: (2026)
Read and Think: An Efficient Step-wise Multimodal Language Model for Document Understanding and Reasoning
von: Zhang, Jinxu
Veröffentlicht: (2024)
von: Zhang, Jinxu
Veröffentlicht: (2024)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
Character-based Outfit Generation with Vision-augmented Style Extraction via LLMs
von: Forouzandehmehr, Najmeh, et al.
Veröffentlicht: (2024)
von: Forouzandehmehr, Najmeh, et al.
Veröffentlicht: (2024)
Low-Data Classification of Historical Music Manuscripts: A Few-Shot Learning Approach
von: Shatri, Elona, et al.
Veröffentlicht: (2024)
von: Shatri, Elona, et al.
Veröffentlicht: (2024)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
Leveraging Customer Feedback for Multi-modal Insight Extraction
von: Mukku, Sandeep Sricharan, et al.
Veröffentlicht: (2024)
von: Mukku, Sandeep Sricharan, et al.
Veröffentlicht: (2024)
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
von: Guo, Yuanhe, et al.
Veröffentlicht: (2025)
von: Guo, Yuanhe, et al.
Veröffentlicht: (2025)
Are They the Same Picture? Adapting Concept Bottleneck Models for Human-AI Collaboration in Image Retrieval
von: Balloli, Vaibhav, et al.
Veröffentlicht: (2024)
von: Balloli, Vaibhav, et al.
Veröffentlicht: (2024)
Compressible and Searchable: AI-native Multi-Modal Retrieval System with Learned Image Compression
von: Luo, Jixiang
Veröffentlicht: (2024)
von: Luo, Jixiang
Veröffentlicht: (2024)
Accelerating Flood Warnings by 10 Hours: The Power of River Network Topology in AI-enhanced Flood Forecasting
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
Retrieval-Guided Generation for Safer Histopathology Image Captioning
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval
von: Liu, Delong, et al.
Veröffentlicht: (2023)
von: Liu, Delong, et al.
Veröffentlicht: (2023)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Very Efficient Listwise Multimodal Reranking for Long Documents
von: Sun, Yiqun, et al.
Veröffentlicht: (2026)
von: Sun, Yiqun, et al.
Veröffentlicht: (2026)
EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection
von: Tang, Wenxin, et al.
Veröffentlicht: (2026)
von: Tang, Wenxin, et al.
Veröffentlicht: (2026)
CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models
von: Alam, Hasan Md Tusfiqur, et al.
Veröffentlicht: (2025)
von: Alam, Hasan Md Tusfiqur, et al.
Veröffentlicht: (2025)
Naiad: Novel Agentic Intelligent Autonomous System for Inland Water Monitoring
von: Baltzi, Eirini, et al.
Veröffentlicht: (2025)
von: Baltzi, Eirini, et al.
Veröffentlicht: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
LookSync: Large-Scale Visual Product Search System for AI-Generated Fashion Looks
von: M, Pradeep, et al.
Veröffentlicht: (2025)
von: M, Pradeep, et al.
Veröffentlicht: (2025)
A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy
von: Ghatwary, Noha, et al.
Veröffentlicht: (2026)
von: Ghatwary, Noha, et al.
Veröffentlicht: (2026)
Hespi: A pipeline for automatically detecting information from hebarium specimen sheets
von: Turnbull, Robert, et al.
Veröffentlicht: (2024)
von: Turnbull, Robert, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
von: Horn, Pius, et al.
Veröffentlicht: (2025) -
From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline
von: Shoilee, Sarah Binta Alam, et al.
Veröffentlicht: (2026) -
Beyond String Matching: Semantic Evaluation of PDF Table Extraction
von: Horn, Pius, et al.
Veröffentlicht: (2026) -
AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
von: Patel, Piyushkumar
Veröffentlicht: (2025) -
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)