Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Yue, Chongjian, Xu, Xinrun, Ma, Xiaojun, Du, Lun, Ding, Zhiming, Han, Shi, Zhang, Dongmei, Zhang, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs
by: Yue, Chongjian, et al.
Published: (2023)
by: Yue, Chongjian, et al.
Published: (2023)
MindRef: Mimicking Human Memory for Hierarchical Reference Retrieval with Fine-Grained Location Awareness
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
TAP4LLM: Table Provider on Sampling, Augmenting, and Packing Semi-structured Data for Large Language Model Reasoning
by: Sui, Yuan, et al.
Published: (2023)
by: Sui, Yuan, et al.
Published: (2023)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
Learning to Generate and Extract: A Multi-Agent Collaboration Framework For Zero-shot Document-level Event Arguments Extraction
by: Zhang, Guangjun, et al.
Published: (2026)
by: Zhang, Guangjun, et al.
Published: (2026)
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
Vulnerability of Text-to-Image Models to Prompt Template Stealing: A Differential Evolution Approach
by: Wu, Yurong, et al.
Published: (2025)
by: Wu, Yurong, et al.
Published: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
by: Liu, Hanbing, et al.
Published: (2025)
by: Liu, Hanbing, et al.
Published: (2025)
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning
by: Guo, Jiaxing, et al.
Published: (2025)
by: Guo, Jiaxing, et al.
Published: (2025)
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)
by: Chai, Mingxu, et al.
Published: (2024)
SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
StrucSum: Graph-Structured Reasoning for Long Document Extractive Summarization with LLMs
by: Yuan, Haohan, et al.
Published: (2025)
by: Yuan, Haohan, et al.
Published: (2025)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
CAST: Achieving Stable LLM-based Text Analysis for Data Analytics
by: Xie, Jinxiang, et al.
Published: (2026)
by: Xie, Jinxiang, et al.
Published: (2026)
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model
by: Gao, Shengxiang, et al.
Published: (2024)
by: Gao, Shengxiang, et al.
Published: (2024)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
by: Hu, Lingxiang, et al.
Published: (2025)
by: Hu, Lingxiang, et al.
Published: (2025)
TwT: Thinking without Tokens by Habitual Reasoning Distillation with Multi-Teachers' Guidance
by: Xu, Jingxian, et al.
Published: (2025)
by: Xu, Jingxian, et al.
Published: (2025)
StructMem: Structured Memory for Long-Horizon Behavior in LLMs
by: Xu, Buqiang, et al.
Published: (2026)
by: Xu, Buqiang, et al.
Published: (2026)
SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs
by: Zhang, Zhenliang, et al.
Published: (2025)
by: Zhang, Zhenliang, et al.
Published: (2025)
LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs via Long Input Fine-Tuning
by: Mao, Yansheng, et al.
Published: (2025)
by: Mao, Yansheng, et al.
Published: (2025)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
Risk Assessment Framework for Code LLMs via Leveraging Internal States
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
by: Shi, Yongxin, et al.
Published: (2025)
by: Shi, Yongxin, et al.
Published: (2025)
Building a Silver-Standard Dataset from NICE Guidelines for Clinical LLMs
by: Ding, Qing, et al.
Published: (2025)
by: Ding, Qing, et al.
Published: (2025)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024)
by: Tan, Weihao, et al.
Published: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
KARPA: A Training-free Method of Adapting Knowledge Graph as References for Large Language Model's Reasoning Path Aggregation
by: Fang, Siyuan, et al.
Published: (2024)
by: Fang, Siyuan, et al.
Published: (2024)
Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs
by: Xiao, Changrong, et al.
Published: (2024)
by: Xiao, Changrong, et al.
Published: (2024)
Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
by: Zhang, Junhao, et al.
Published: (2025)
by: Zhang, Junhao, et al.
Published: (2025)
Tackling Long Code Search with Splitting, Encoding, and Aggregating
by: Hu, Fan, et al.
Published: (2022)
by: Hu, Fan, et al.
Published: (2022)
TAROT: A Hierarchical Framework with Multitask Co-Pretraining on Semi-Structured Data towards Effective Person-Job Fit
by: Cao, Yihan, et al.
Published: (2024)
by: Cao, Yihan, et al.
Published: (2024)
Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection
by: Zhang, Huixuan, et al.
Published: (2023)
by: Zhang, Huixuan, et al.
Published: (2023)
Less is More for Long Document Summary Evaluation by LLMs
by: Wu, Yunshu, et al.
Published: (2023)
by: Wu, Yunshu, et al.
Published: (2023)
SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs
by: Sun, Jie, et al.
Published: (2026)
by: Sun, Jie, et al.
Published: (2026)
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
by: Wu, Kaifeng, et al.
Published: (2025)
by: Wu, Kaifeng, et al.
Published: (2025)
Similar Items
-
Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs
by: Yue, Chongjian, et al.
Published: (2023) -
MindRef: Mimicking Human Memory for Hierarchical Reference Retrieval with Fine-Grained Location Awareness
by: Wang, Ye, et al.
Published: (2024) -
TAP4LLM: Table Provider on Sampling, Augmenting, and Packing Semi-structured Data for Large Language Model Reasoning
by: Sui, Yuan, et al.
Published: (2023) -
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
by: Hu, Mengkang, et al.
Published: (2024) -
Learning to Generate and Extract: A Multi-Agent Collaboration Framework For Zero-shot Document-level Event Arguments Extraction
by: Zhang, Guangjun, et al.
Published: (2026)