READoc: A Unified Benchmark for Realistic Document Structured Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zichao, Abulaiti, Aizier, Lu, Yaojie, Chen, Xuanang, Zheng, Jia, Lin, Hongyu, Han, Xianpei, Sun, Le |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seg2Act: Global Context-aware Action Generation for Document Logical Structuring
von: Li, Zichao, et al.
Veröffentlicht: (2024)
von: Li, Zichao, et al.
Veröffentlicht: (2024)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction
von: Zhong, Tianyun, et al.
Veröffentlicht: (2025)
von: Zhong, Tianyun, et al.
Veröffentlicht: (2025)
Open Grounded Planning: Challenges and Benchmark Construction
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?
von: Bian, Ning, et al.
Veröffentlicht: (2024)
von: Bian, Ning, et al.
Veröffentlicht: (2024)
REInstruct: Building Instruction Data from Unlabeled Corpus
von: Chen, Shu, et al.
Veröffentlicht: (2024)
von: Chen, Shu, et al.
Veröffentlicht: (2024)
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2024)
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2024)
StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
Multi-Facet Counterfactual Learning for Content Quality Evaluation
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
AI-Salesman: Towards Reliable Large Language Model Driven Telemarketing
von: Zhang, Qingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Qingyu, et al.
Veröffentlicht: (2025)
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
Few-shot Named Entity Recognition via Superposition Concept Discrimination
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
von: Mo, Guozhao, et al.
Veröffentlicht: (2025)
von: Mo, Guozhao, et al.
Veröffentlicht: (2025)
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
von: Peng, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Peng, Xiaoxuan, et al.
Veröffentlicht: (2026)
SoFA: Shielded On-the-fly Alignment via Priority Rule Following
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
von: Su, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Su, Xiaoyan, et al.
Veröffentlicht: (2026)
RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2025)
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2025)
ChatGPT is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language Models
von: Bian, Ning, et al.
Veröffentlicht: (2023)
von: Bian, Ning, et al.
Veröffentlicht: (2023)
Robustness of Structured Data Extraction from Perspectively Distorted Documents
von: Nakada, Hyakka, et al.
Veröffentlicht: (2025)
von: Nakada, Hyakka, et al.
Veröffentlicht: (2025)
The Life Cycle of Knowledge in Big Language Models: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2023)
von: Cao, Boxi, et al.
Veröffentlicht: (2023)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
von: Zheng, Hao, et al.
Veröffentlicht: (2025)
von: Zheng, Hao, et al.
Veröffentlicht: (2025)
Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching
von: Wang, Tianshu, et al.
Veröffentlicht: (2024)
von: Wang, Tianshu, et al.
Veröffentlicht: (2024)
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
von: Li, Zichao, et al.
Veröffentlicht: (2025)
von: Li, Zichao, et al.
Veröffentlicht: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
von: Sun, Hao, et al.
Veröffentlicht: (2026)
von: Sun, Hao, et al.
Veröffentlicht: (2026)
UEval: A Benchmark for Unified Multimodal Generation
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
MetaphorVU: Towards Metaphorical Video Understanding
von: Li, Zhuoqun, et al.
Veröffentlicht: (2026)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2026)
Academically intelligent LLMs are not necessarily socially intelligent
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seg2Act: Global Context-aware Action Generation for Document Logical Structuring
von: Li, Zichao, et al.
Veröffentlicht: (2024) -
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025) -
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025) -
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025) -
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)