ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
Fuente:
arXiv
Guardado en:
| Autores principales: | Sibue, Mathieu, Garza, Andres Muñoz, Mensah, Samuel, Shetty, Pranav, Ma, Zhiqiang, Liu, Xiaomo, Veloso, Manuela |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
por: Zmigrod, Ran, et al.
Publicado: (2024)
por: Zmigrod, Ran, et al.
Publicado: (2024)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
por: Shetty, Pranav, et al.
Publicado: (2025)
por: Shetty, Pranav, et al.
Publicado: (2025)
Detecting Non-Membership in LLM Training Data via Rank Correlations
por: Shetty, Pranav, et al.
Publicado: (2026)
por: Shetty, Pranav, et al.
Publicado: (2026)
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
por: Zmigrod, Ran, et al.
Publicado: (2024)
por: Zmigrod, Ran, et al.
Publicado: (2024)
DocLLM: A layout-aware generative language model for multimodal document understanding
por: Wang, Dongsheng, et al.
Publicado: (2023)
por: Wang, Dongsheng, et al.
Publicado: (2023)
TASER: Table Agents for Schema-guided Extraction and Recommendation
por: Cho, Nicole, et al.
Publicado: (2025)
por: Cho, Nicole, et al.
Publicado: (2025)
Meta-RAG on Large Codebases Using Code Summarization
por: Tawosi, Vali, et al.
Publicado: (2025)
por: Tawosi, Vali, et al.
Publicado: (2025)
LLM Agents for Automated Dependency Upgrades
por: Tawosi, Vali, et al.
Publicado: (2025)
por: Tawosi, Vali, et al.
Publicado: (2025)
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
por: S, Santosh T. Y. S., et al.
Publicado: (2025)
por: S, Santosh T. Y. S., et al.
Publicado: (2025)
StrucSum: Graph-Structured Reasoning for Long Document Extractive Summarization with LLMs
por: Yuan, Haohan, et al.
Publicado: (2025)
por: Yuan, Haohan, et al.
Publicado: (2025)
Distill and Align Decomposition for Enhanced Claim Verification
por: Magomere, Jabez, et al.
Publicado: (2026)
por: Magomere, Jabez, et al.
Publicado: (2026)
Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
por: Liu, Shikun, et al.
Publicado: (2025)
por: Liu, Shikun, et al.
Publicado: (2025)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
por: Haque, Mirazul, et al.
Publicado: (2026)
por: Haque, Mirazul, et al.
Publicado: (2026)
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
por: Tang, Xiangru, et al.
Publicado: (2023)
por: Tang, Xiangru, et al.
Publicado: (2023)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
por: Li, Xianzhi, et al.
Publicado: (2024)
por: Li, Xianzhi, et al.
Publicado: (2024)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
por: Lyu, Pengyuan, et al.
Publicado: (2024)
por: Lyu, Pengyuan, et al.
Publicado: (2024)
StrucADT: Generating Structure-controlled 3D Point Clouds with Adjacency Diffusion Transformer
por: Shu, Zhenyu, et al.
Publicado: (2025)
por: Shu, Zhenyu, et al.
Publicado: (2025)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
por: Gu, Zhouhong, et al.
Publicado: (2024)
por: Gu, Zhouhong, et al.
Publicado: (2024)
BEMEval-Doc2Schema: Benchmarking Large Language Models for Structured Data Extraction in Building Energy Modeling
por: Jia, Yiyuan, et al.
Publicado: (2026)
por: Jia, Yiyuan, et al.
Publicado: (2026)
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
por: Yin, Jiaqi, et al.
Publicado: (2025)
por: Yin, Jiaqi, et al.
Publicado: (2025)
Common Foundations for SHACL, ShEx, and PG-Schema
por: Ahmetaj, S., et al.
Publicado: (2025)
por: Ahmetaj, S., et al.
Publicado: (2025)
SchemaCoder: Automatic Log Schema Extraction Coder with Residual Q-Tree Boosting
por: Wan, Lily Jiaxin, et al.
Publicado: (2025)
por: Wan, Lily Jiaxin, et al.
Publicado: (2025)
READoc: A Unified Benchmark for Realistic Document Structured Extraction
por: Li, Zichao, et al.
Publicado: (2024)
por: Li, Zichao, et al.
Publicado: (2024)
Schema as Parameterized Tools for Universal Information Extraction
por: Liang, Sheng, et al.
Publicado: (2025)
por: Liang, Sheng, et al.
Publicado: (2025)
Intelligent Execution through Plan Analysis
por: Borrajo, Daniel, et al.
Publicado: (2024)
por: Borrajo, Daniel, et al.
Publicado: (2024)
DocGraphLM: Documental Graph Language Model for Information Extraction
por: Wang, Dongsheng, et al.
Publicado: (2024)
por: Wang, Dongsheng, et al.
Publicado: (2024)
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
por: Barzelay, Udi, et al.
Publicado: (2026)
por: Barzelay, Udi, et al.
Publicado: (2026)
Schema-Driven Information Extraction from Heterogeneous Tables
por: Bai, Fan, et al.
Publicado: (2023)
por: Bai, Fan, et al.
Publicado: (2023)
OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System
por: Luo, Yujie, et al.
Publicado: (2024)
por: Luo, Yujie, et al.
Publicado: (2024)
Artificial Intelligence Applications in Environmental Monitoring
por: Samuel Mensah, SM
Publicado: (2026)
por: Samuel Mensah, SM
Publicado: (2026)
From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
por: Petrov, Alex, et al.
Publicado: (2026)
por: Petrov, Alex, et al.
Publicado: (2026)
Where is this coming from? Making groundedness count in the evaluation of Document VQA models
por: Nourbakhsh, Armineh, et al.
Publicado: (2025)
por: Nourbakhsh, Armineh, et al.
Publicado: (2025)
Benchmarking Table Extraction from Heterogeneous Scientific Extraction Documents
por: Soric, Marijan, et al.
Publicado: (2025)
por: Soric, Marijan, et al.
Publicado: (2025)
Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation
por: Liang, Sheng, et al.
Publicado: (2025)
por: Liang, Sheng, et al.
Publicado: (2025)
Schema-Grounded LLM Extraction for FHIR Patient Digital Twins
por: Brens, Rafael, et al.
Publicado: (2026)
por: Brens, Rafael, et al.
Publicado: (2026)
PARSE: LLM Driven Schema Optimization for Reliable Entity Extraction
por: Shrimal, Anubhav, et al.
Publicado: (2025)
por: Shrimal, Anubhav, et al.
Publicado: (2025)
Struc2mapGAN: improving synthetic cryo-EM density maps with generative adversarial networks
por: Zhang, Chenwei, et al.
Publicado: (2024)
por: Zhang, Chenwei, et al.
Publicado: (2024)
GDC-SM: The GDC Schema Matching Benchmark
por: Santos, Aécio, et al.
Publicado: (2025)
por: Santos, Aécio, et al.
Publicado: (2025)
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
por: Roberts, Josselin Somerville, et al.
Publicado: (2024)
por: Roberts, Josselin Somerville, et al.
Publicado: (2024)
Robust Detection of Synthetic Tabular Data under Schema Variability
por: Kindji, G. Charbel N., et al.
Publicado: (2025)
por: Kindji, G. Charbel N., et al.
Publicado: (2025)
Ejemplares similares
-
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
por: Zmigrod, Ran, et al.
Publicado: (2024) -
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
por: Shetty, Pranav, et al.
Publicado: (2025) -
Detecting Non-Membership in LLM Training Data via Rank Correlations
por: Shetty, Pranav, et al.
Publicado: (2026) -
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
por: Zmigrod, Ran, et al.
Publicado: (2024) -
DocLLM: A layout-aware generative language model for multimodal document understanding
por: Wang, Dongsheng, et al.
Publicado: (2023)