"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zmigrod, Ran, Shetty, Pranav, Sibue, Mathieu, Ma, Zhiqiang, Nourbakhsh, Armineh, Liu, Xiaomo, Veloso, Manuela |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
by: Sibue, Mathieu, et al.
Published: (2026)
by: Sibue, Mathieu, et al.
Published: (2026)
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
TreeForm: End-to-end Annotation and Evaluation for Form Document Parsing
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)
by: Wang, Dongsheng, et al.
Published: (2023)
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025)
by: S, Santosh T. Y. S., et al.
Published: (2025)
DocGraphLM: Documental Graph Language Model for Information Extraction
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
by: Shetty, Pranav, et al.
Published: (2025)
by: Shetty, Pranav, et al.
Published: (2025)
Detecting Non-Membership in LLM Training Data via Rank Correlations
by: Shetty, Pranav, et al.
Published: (2026)
by: Shetty, Pranav, et al.
Published: (2026)
Translating between SQL Dialects for Cloud Migration
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
Where is this coming from? Making groundedness count in the evaluation of Document VQA models
by: Nourbakhsh, Armineh, et al.
Published: (2025)
by: Nourbakhsh, Armineh, et al.
Published: (2025)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
by: Li, Xianzhi, et al.
Published: (2024)
by: Li, Xianzhi, et al.
Published: (2024)
Belief and Persuasion in Scientific Discourse on Social Media: A Study of the COVID-19 Pandemic
by: Alamir, Salwa, et al.
Published: (2024)
by: Alamir, Salwa, et al.
Published: (2024)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
by: Shen, Jiyuan, et al.
Published: (2026)
by: Shen, Jiyuan, et al.
Published: (2026)
Log Summarisation for Defect Evolution Analysis
by: Dolga, Rares, et al.
Published: (2024)
by: Dolga, Rares, et al.
Published: (2024)
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law
by: Chen, Zhiyu Zoey, et al.
Published: (2024)
by: Chen, Zhiyu Zoey, et al.
Published: (2024)
Information Extraction From Fiscal Documents Using LLMs
by: Aggarwal, Vikram, et al.
Published: (2025)
by: Aggarwal, Vikram, et al.
Published: (2025)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
by: Watson, William, et al.
Published: (2026)
by: Watson, William, et al.
Published: (2026)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
by: Ma, Youmi, et al.
Published: (2024)
by: Ma, Youmi, et al.
Published: (2024)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction
by: Liu, Jiang, et al.
Published: (2024)
by: Liu, Jiang, et al.
Published: (2024)
TASER: Table Agents for Schema-guided Extraction and Recommendation
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
by: Yue, Chongjian, et al.
Published: (2024)
by: Yue, Chongjian, et al.
Published: (2024)
SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
by: Tan, Nelvin, et al.
Published: (2026)
by: Tan, Nelvin, et al.
Published: (2026)
Problem Solved? Information Extraction Design Space for Layout-Rich Documents using LLMs
by: Colakoglu, Gaye, et al.
Published: (2025)
by: Colakoglu, Gaye, et al.
Published: (2025)
Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction
by: Qi, Ji, et al.
Published: (2023)
by: Qi, Ji, et al.
Published: (2023)
Neurosymbolic Information Extraction from Transactional Documents
by: Hemmer, Arthur, et al.
Published: (2025)
by: Hemmer, Arthur, et al.
Published: (2025)
SLIDE: Sliding Localized Information for Document Extraction
by: Singh, Divyansh, et al.
Published: (2025)
by: Singh, Divyansh, et al.
Published: (2025)
Quantitative Information Extraction from Humanitarian Documents
by: Liberatore, Daniele, et al.
Published: (2024)
by: Liberatore, Daniele, et al.
Published: (2024)
Rethinking the Role of LLMs in Time Series Forecasting
by: Qiu, Xin, et al.
Published: (2026)
by: Qiu, Xin, et al.
Published: (2026)
FlowMind: Automatic Workflow Generation with LLMs
by: Zeng, Zhen, et al.
Published: (2024)
by: Zeng, Zhen, et al.
Published: (2024)
KIEval: Evaluation Metric for Document Key Information Extraction
by: Khang, Minsoo, et al.
Published: (2025)
by: Khang, Minsoo, et al.
Published: (2025)
Enriching Datasets with Demographics through Large Language Models: What's in a Name?
by: AlNuaimi, Khaled, et al.
Published: (2024)
by: AlNuaimi, Khaled, et al.
Published: (2024)
Enhancing Document-level Argument Extraction with Definition-augmented Heuristic-driven Prompting for LLMs
by: Sun, Tongyue, et al.
Published: (2024)
by: Sun, Tongyue, et al.
Published: (2024)
Research on Information Extraction of LCSTS Dataset Based on an Improved BERTSum-LSTM Model
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
Low-resource Information Extraction with the European Clinical Case Corpus
by: Ghosh, Soumitra, et al.
Published: (2025)
by: Ghosh, Soumitra, et al.
Published: (2025)
Beyond path selection: Better LLMs for Scientific Information Extraction with MimicSFT and Relevance and Rule-induced(R$^2$)GRPO
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
Graph-Augmented Relation Extraction Model with LLMs-Generated Support Document
by: Dong, Vicky, et al.
Published: (2024)
by: Dong, Vicky, et al.
Published: (2024)
CodeMirage: Hallucinations in Code Generated by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
Similar Items
-
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
by: Sibue, Mathieu, et al.
Published: (2026) -
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
by: Zmigrod, Ran, et al.
Published: (2024) -
TreeForm: End-to-end Annotation and Evaluation for Form Document Parsing
by: Zmigrod, Ran, et al.
Published: (2024) -
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023) -
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025)