Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Saurabh K., Raj, Sachin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unified Multimodal Interleaved Document Representation for Retrieval
por: Lee, Jaewoo, et al.
Publicado: (2024)
por: Lee, Jaewoo, et al.
Publicado: (2024)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
por: Opsahl-Ong, Krista, et al.
Publicado: (2026)
por: Opsahl-Ong, Krista, et al.
Publicado: (2026)
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
por: Wu, Hui, et al.
Publicado: (2026)
por: Wu, Hui, et al.
Publicado: (2026)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
por: Khanghah, Kiarash Naghavi, et al.
Publicado: (2026)
por: Khanghah, Kiarash Naghavi, et al.
Publicado: (2026)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
por: Thakur, Nandan, et al.
Publicado: (2025)
por: Thakur, Nandan, et al.
Publicado: (2025)
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
por: Su, Weihang, et al.
Publicado: (2025)
por: Su, Weihang, et al.
Publicado: (2025)
Transformers Remember First, Forget Last: Dual-Process Interference in LLMs
por: Chattaraj, Sourav, et al.
Publicado: (2026)
por: Chattaraj, Sourav, et al.
Publicado: (2026)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
por: Tan, Jiejun, et al.
Publicado: (2025)
por: Tan, Jiejun, et al.
Publicado: (2025)
The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
por: Huang, Chau-Jian
Publicado: (2024)
por: Huang, Chau-Jian
Publicado: (2024)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
por: Killingback, Julian, et al.
Publicado: (2025)
por: Killingback, Julian, et al.
Publicado: (2025)
Structured Attention Matters to Multimodal LLMs in Document Understanding
por: Liu, Chang, et al.
Publicado: (2025)
por: Liu, Chang, et al.
Publicado: (2025)
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
por: Cui, Xuanming, et al.
Publicado: (2025)
por: Cui, Xuanming, et al.
Publicado: (2025)
SwasthLLM: a Unified Cross-Lingual, Multi-Task, and Meta-Learning Zero-Shot Framework for Medical Diagnosis Using Contrastive Representations
por: Sar, Ayan, et al.
Publicado: (2025)
por: Sar, Ayan, et al.
Publicado: (2025)
QAEA-DR: A Unified Text Augmentation Framework for Dense Retrieval
por: Tan, Hongming, et al.
Publicado: (2024)
por: Tan, Hongming, et al.
Publicado: (2024)
Model-Document Protocol for AI Search
por: Qian, Hongjin, et al.
Publicado: (2025)
por: Qian, Hongjin, et al.
Publicado: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
por: Islamaj, Rezarta, et al.
Publicado: (2026)
por: Islamaj, Rezarta, et al.
Publicado: (2026)
TURA: Tool-Augmented Unified Retrieval Agent for AI Search
por: Zhao, Zhejun, et al.
Publicado: (2025)
por: Zhao, Zhejun, et al.
Publicado: (2025)
A Semi-supervised Scalable Unified Framework for E-commerce Query Classification
por: Yuan, Chunyuan, et al.
Publicado: (2025)
por: Yuan, Chunyuan, et al.
Publicado: (2025)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
por: Kim, Eric Y., et al.
Publicado: (2026)
por: Kim, Eric Y., et al.
Publicado: (2026)
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
por: Su, Weihang, et al.
Publicado: (2025)
por: Su, Weihang, et al.
Publicado: (2025)
Towards Personalized Deep Research: Benchmarks and Evaluations
por: Liang, Yuan, et al.
Publicado: (2025)
por: Liang, Yuan, et al.
Publicado: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
por: Kuissi, Nathan, et al.
Publicado: (2026)
por: Kuissi, Nathan, et al.
Publicado: (2026)
On the impact of retrieved content representations in RAG Pipelines
por: Ross, Jonathan J, et al.
Publicado: (2026)
por: Ross, Jonathan J, et al.
Publicado: (2026)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
por: Wu, Yutao, et al.
Publicado: (2025)
por: Wu, Yutao, et al.
Publicado: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
por: Dong, Kuicai, et al.
Publicado: (2025)
por: Dong, Kuicai, et al.
Publicado: (2025)
Bias-Aware Agent: Enhancing Fairness in AI-Driven Knowledge Retrieval
por: Singh, Karanbir, et al.
Publicado: (2025)
por: Singh, Karanbir, et al.
Publicado: (2025)
Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components
por: Alsubhi, Jumana, et al.
Publicado: (2025)
por: Alsubhi, Jumana, et al.
Publicado: (2025)
A Semantic Search Pipeline for Causality-driven Adhoc Information Retrieval
por: Dalal, Dhairya, et al.
Publicado: (2025)
por: Dalal, Dhairya, et al.
Publicado: (2025)
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
por: Wu, Kaifeng, et al.
Publicado: (2025)
por: Wu, Kaifeng, et al.
Publicado: (2025)
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
por: Kang, Yue, et al.
Publicado: (2026)
por: Kang, Yue, et al.
Publicado: (2026)
Enhancing Question Answering for Enterprise Knowledge Bases using Large Language Models
por: Jiang, Feihu, et al.
Publicado: (2024)
por: Jiang, Feihu, et al.
Publicado: (2024)
The Text Classification Pipeline: Starting Shallow going Deeper
por: Siino, Marco, et al.
Publicado: (2024)
por: Siino, Marco, et al.
Publicado: (2024)
Docs2KG: Unified Knowledge Graph Construction from Heterogeneous Documents Assisted by Large Language Models
por: Sun, Qiang, et al.
Publicado: (2024)
por: Sun, Qiang, et al.
Publicado: (2024)
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
por: Merola, Carlo, et al.
Publicado: (2025)
por: Merola, Carlo, et al.
Publicado: (2025)
EnterpriseEM: Fine-tuned Embeddings for Enterprise Semantic Search
por: Rathinasamy, Kamalkumar, et al.
Publicado: (2024)
por: Rathinasamy, Kamalkumar, et al.
Publicado: (2024)
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
por: Alawwad, Hessa A., et al.
Publicado: (2025)
por: Alawwad, Hessa A., et al.
Publicado: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
por: Markowitz, Elan, et al.
Publicado: (2025)
por: Markowitz, Elan, et al.
Publicado: (2025)
A Framework for Leveraging Partially-Labeled Data for Product Attribute-Value Identification
por: Subhalingam, D., et al.
Publicado: (2024)
por: Subhalingam, D., et al.
Publicado: (2024)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
por: Li, Zhuofeng, et al.
Publicado: (2026)
por: Li, Zhuofeng, et al.
Publicado: (2026)
Improving Robustness of Tabular Retrieval via Representational Stability
por: Bhandari, Kushal Raj, et al.
Publicado: (2026)
por: Bhandari, Kushal Raj, et al.
Publicado: (2026)
Ejemplares similares
-
Unified Multimodal Interleaved Document Representation for Retrieval
por: Lee, Jaewoo, et al.
Publicado: (2024) -
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
por: Opsahl-Ong, Krista, et al.
Publicado: (2026) -
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
por: Wu, Hui, et al.
Publicado: (2026) -
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
por: Khanghah, Kiarash Naghavi, et al.
Publicado: (2026) -
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
por: Thakur, Nandan, et al.
Publicado: (2025)