Visually Guided Generative Text-Layout Pre-training for Document Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Zhiming, Bai, Haoli, Hou, Lu, Wei, Jiansheng, Jiang, Xin, Liu, Qun, Wong, Kam-Fai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling the Deficiencies of Pre-trained Text-and-Layout Models in Real-world Visually-rich Document Information Extraction
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
by: Wan, Zhongwei, et al.
Published: (2022)
by: Wan, Zhongwei, et al.
Published: (2022)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
by: Kwan, Wai-Chung, et al.
Published: (2024)
by: Kwan, Wai-Chung, et al.
Published: (2024)
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
by: Lin, Luyang, et al.
Published: (2025)
by: Lin, Luyang, et al.
Published: (2025)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
by: Wang, Zezhong, et al.
Published: (2024)
by: Wang, Zezhong, et al.
Published: (2024)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
by: Cui, Wenqian, et al.
Published: (2025)
by: Cui, Wenqian, et al.
Published: (2025)
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
by: Wang, Zezhong, et al.
Published: (2025)
by: Wang, Zezhong, et al.
Published: (2025)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
by: Wang, Zezhong, et al.
Published: (2024)
by: Wang, Zezhong, et al.
Published: (2024)
DocPolarBERT: A Pre-trained Model for Document Understanding with Relative Polar Coordinate Encoding of Layout Structures
by: Uthayasooriyar, Benno, et al.
Published: (2025)
by: Uthayasooriyar, Benno, et al.
Published: (2025)
Text-to-Code Generation with Modality-relative Pre-training
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
by: Li, Kaican, et al.
Published: (2025)
by: Li, Kaican, et al.
Published: (2025)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
by: Kwan, Wai-Chung, et al.
Published: (2023)
by: Kwan, Wai-Chung, et al.
Published: (2023)
WHERE and WHICH: Iterative Debate for Biomedical Synthetic Data Augmentation
by: Zhao, Zhengyi, et al.
Published: (2025)
by: Zhao, Zhengyi, et al.
Published: (2025)
IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators
by: Lin, Luyang, et al.
Published: (2024)
by: Lin, Luyang, et al.
Published: (2024)
Text Embeddings by Weakly-Supervised Contrastive Pre-training
by: Wang, Liang, et al.
Published: (2022)
by: Wang, Liang, et al.
Published: (2022)
DataVisT5: A Pre-trained Language Model for Jointly Understanding Text and Data Visualization
by: Wan, Zhuoyue, et al.
Published: (2024)
by: Wan, Zhuoyue, et al.
Published: (2024)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
by: He, Ziwei, et al.
Published: (2025)
by: He, Ziwei, et al.
Published: (2025)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
by: Yuan, Jian, et al.
Published: (2025)
by: Yuan, Jian, et al.
Published: (2025)
Pre-training Distillation for Large Language Models: A Design Space Exploration
by: Peng, Hao, et al.
Published: (2024)
by: Peng, Hao, et al.
Published: (2024)
Dual-Density Inference for Efficient Language Model Reasoning
by: Zhao, Zhengyi, et al.
Published: (2025)
by: Zhao, Zhengyi, et al.
Published: (2025)
FReM: A Flexible Reasoning Mechanism for Balancing Quick and Slow Thinking in Long-Context Question Answering
by: Zhao, Zhengyi, et al.
Published: (2025)
by: Zhao, Zhengyi, et al.
Published: (2025)
ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
A Survey of Pre-trained Language Models for Processing Scientific Text
by: Ho, Xanh, et al.
Published: (2024)
by: Ho, Xanh, et al.
Published: (2024)
Development of Cognitive Intelligence in Pre-trained Language Models
by: Shah, Raj Sanjay, et al.
Published: (2024)
by: Shah, Raj Sanjay, et al.
Published: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
RecGPT: Generative Pre-training for Text-based Recommendation
by: Ngo, Hoang, et al.
Published: (2024)
by: Ngo, Hoang, et al.
Published: (2024)
Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models
by: Wang, Lingzhi, et al.
Published: (2024)
by: Wang, Lingzhi, et al.
Published: (2024)
LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding
by: Sourati, Zhivar, et al.
Published: (2025)
by: Sourati, Zhivar, et al.
Published: (2025)
Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
by: Sankar, Sanjana, et al.
Published: (2025)
by: Sankar, Sanjana, et al.
Published: (2025)
Faster and Better LLMs via Latency-Aware Test-Time Scaling
by: Wang, Zili, et al.
Published: (2025)
by: Wang, Zili, et al.
Published: (2025)
E$^3$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
by: Yuan, Tao, et al.
Published: (2025)
by: Yuan, Tao, et al.
Published: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
by: Wang, Zezhong, et al.
Published: (2023)
by: Wang, Zezhong, et al.
Published: (2023)
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models
by: Zhang, Ruiqi, et al.
Published: (2026)
by: Zhang, Ruiqi, et al.
Published: (2026)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
Similar Items
-
Unveiling the Deficiencies of Pre-trained Text-and-Layout Models in Real-world Visually-rich Document Information Extraction
by: Zhang, Chong, et al.
Published: (2024) -
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
by: Wan, Zhongwei, et al.
Published: (2022) -
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
by: Chen, Liang, et al.
Published: (2025) -
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
by: Kwan, Wai-Chung, et al.
Published: (2024) -
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
by: Lin, Luyang, et al.
Published: (2025)