Beyond String Matching: Semantic Evaluation of PDF Table Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Horn, Pius, Keuper, Janis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
von: Horn, Pius, et al.
Veröffentlicht: (2025)
von: Horn, Pius, et al.
Veröffentlicht: (2025)
TalentMine: LLM-Based Extraction and Question-Answering from Multimodal Talent Tables
von: Mannam, Varun, et al.
Veröffentlicht: (2025)
von: Mannam, Varun, et al.
Veröffentlicht: (2025)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
von: Xu, Tianyi, et al.
Veröffentlicht: (2026)
von: Xu, Tianyi, et al.
Veröffentlicht: (2026)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
von: Barrios, Wayner, et al.
Veröffentlicht: (2026)
von: Barrios, Wayner, et al.
Veröffentlicht: (2026)
NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence
von: Ulla, Aman
Veröffentlicht: (2026)
von: Ulla, Aman
Veröffentlicht: (2026)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
von: Askari, Arian, et al.
Veröffentlicht: (2025)
von: Askari, Arian, et al.
Veröffentlicht: (2025)
RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?
von: Ghosh, Arijit, et al.
Veröffentlicht: (2026)
von: Ghosh, Arijit, et al.
Veröffentlicht: (2026)
EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection
von: Tang, Wenxin, et al.
Veröffentlicht: (2026)
von: Tang, Wenxin, et al.
Veröffentlicht: (2026)
RAPTOR: Refined Approach for Product Table Object Recognition
von: Thomas, Eliott, et al.
Veröffentlicht: (2025)
von: Thomas, Eliott, et al.
Veröffentlicht: (2025)
Leveraging Customer Feedback for Multi-modal Insight Extraction
von: Mukku, Sandeep Sricharan, et al.
Veröffentlicht: (2024)
von: Mukku, Sandeep Sricharan, et al.
Veröffentlicht: (2024)
Learning Visual Composition through Improved Semantic Guidance
von: Stone, Austin, et al.
Veröffentlicht: (2024)
von: Stone, Austin, et al.
Veröffentlicht: (2024)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
Character-based Outfit Generation with Vision-augmented Style Extraction via LLMs
von: Forouzandehmehr, Najmeh, et al.
Veröffentlicht: (2024)
von: Forouzandehmehr, Najmeh, et al.
Veröffentlicht: (2024)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
von: Wen, Tiansheng, et al.
Veröffentlicht: (2025)
von: Wen, Tiansheng, et al.
Veröffentlicht: (2025)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
von: Lim, Ho Hung, et al.
Veröffentlicht: (2026)
von: Lim, Ho Hung, et al.
Veröffentlicht: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline
von: Shoilee, Sarah Binta Alam, et al.
Veröffentlicht: (2026)
von: Shoilee, Sarah Binta Alam, et al.
Veröffentlicht: (2026)
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
von: Chen, Lei, et al.
Veröffentlicht: (2026)
von: Chen, Lei, et al.
Veröffentlicht: (2026)
Your Embedding Model is SMARTer Than You Think
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
Efficient Logic Gate Networks for Video Copy Detection
von: Fojcik, Katarzyna
Veröffentlicht: (2026)
von: Fojcik, Katarzyna
Veröffentlicht: (2026)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
Good Scores, Bad Data: A Metric for Multimodal Coherence
von: Srinivasan, Vasundra
Veröffentlicht: (2026)
von: Srinivasan, Vasundra
Veröffentlicht: (2026)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
Multi-task Cross-modal Learning for Chest X-ray Image Retrieval
von: Liang, Zhaohui, et al.
Veröffentlicht: (2026)
von: Liang, Zhaohui, et al.
Veröffentlicht: (2026)
PEARL: Personalized Streaming Video Understanding Model
von: Zheng, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zheng, Yuanhong, et al.
Veröffentlicht: (2026)
A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy
von: Ghatwary, Noha, et al.
Veröffentlicht: (2026)
von: Ghatwary, Noha, et al.
Veröffentlicht: (2026)
CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval
von: Afzal, Zahra Rahimi, et al.
Veröffentlicht: (2026)
von: Afzal, Zahra Rahimi, et al.
Veröffentlicht: (2026)
Retrieval-Guided Generation for Safer Histopathology Image Captioning
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
Open-SAT: LLM-Guided Query Embedding Refinement for Open-Vocabulary Object Retrieval in Satellite Imagery
von: Arefeen, Md Adnan, et al.
Veröffentlicht: (2026)
von: Arefeen, Md Adnan, et al.
Veröffentlicht: (2026)
Benchmark Granularity and Model Robustness for Image-Text Retrieval
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Weakly Supervised Deep Hyperspherical Quantization for Image Retrieval
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
Low-Data Classification of Historical Music Manuscripts: A Few-Shot Learning Approach
von: Shatri, Elona, et al.
Veröffentlicht: (2024)
von: Shatri, Elona, et al.
Veröffentlicht: (2024)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
von: Luo, Enming, et al.
Veröffentlicht: (2024)
von: Luo, Enming, et al.
Veröffentlicht: (2024)
Accelerating Flood Warnings by 10 Hours: The Power of River Network Topology in AI-enhanced Flood Forecasting
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
von: Horn, Pius, et al.
Veröffentlicht: (2025) -
TalentMine: LLM-Based Extraction and Question-Answering from Multimodal Talent Tables
von: Mannam, Varun, et al.
Veröffentlicht: (2025) -
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024) -
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
von: Xu, Tianyi, et al.
Veröffentlicht: (2026) -
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
von: Zhu, Jing, et al.
Veröffentlicht: (2025)