SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Yihao, Han, Soyeon Caren, Jiang, Yanbei, Li, Yan, Li, Zechuan, Peng, Yifan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
by: Xiang, Biao, et al.
Published: (2026)
by: Xiang, Biao, et al.
Published: (2026)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
by: Park, Jiwon, et al.
Published: (2025)
by: Park, Jiwon, et al.
Published: (2025)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Graph Neural Networks for Text Classification: A Survey
by: Wang, Kunze, et al.
Published: (2023)
by: Wang, Kunze, et al.
Published: (2023)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs
by: Pratama, Dhita Putri, et al.
Published: (2026)
by: Pratama, Dhita Putri, et al.
Published: (2026)
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
by: Ding, Yihao, et al.
Published: (2026)
by: Ding, Yihao, et al.
Published: (2026)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
ToolTree: Efficient LLM Agent Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)
by: Akasaka, Muku, et al.
Published: (2026)
UniCast: A Unified Framework for Instance-Conditioned Multimodal Time-Series Forecasting
by: Park, Sehyuk, et al.
Published: (2025)
by: Park, Sehyuk, et al.
Published: (2025)
SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models
by: Rubinstein, Beny, et al.
Published: (2026)
by: Rubinstein, Beny, et al.
Published: (2026)
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024)
by: Wu, Yiheng, et al.
Published: (2024)
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
by: Zhao, Haoran, et al.
Published: (2026)
by: Zhao, Haoran, et al.
Published: (2026)
MSG-Chart: Multimodal Scene Graph for ChartQA
by: Dai, Yue, et al.
Published: (2024)
by: Dai, Yue, et al.
Published: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025)
by: Dai, Yue, et al.
Published: (2025)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
VaeDiff-DocRE: End-to-end Data Augmentation Framework for Document-level Relation Extraction
by: Tran, Khai Phan, et al.
Published: (2024)
by: Tran, Khai Phan, et al.
Published: (2024)
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
by: Li, Haoyi, et al.
Published: (2025)
by: Li, Haoyi, et al.
Published: (2025)
MUSEKG: A Knowledge Graph Over Museum Collections
by: Li, Jinhao, et al.
Published: (2025)
by: Li, Jinhao, et al.
Published: (2025)
EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
by: Yuan, Angela Yifei, et al.
Published: (2025)
by: Yuan, Angela Yifei, et al.
Published: (2025)
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024)
by: Ng, Thye Shan, et al.
Published: (2024)
LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach
by: Pilligua, Maria, et al.
Published: (2024)
by: Pilligua, Maria, et al.
Published: (2024)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
by: Zhao, Zhiyuan, et al.
Published: (2024)
by: Zhao, Zhiyuan, et al.
Published: (2024)
ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
by: Zhang, Hengrui, et al.
Published: (2025)
by: Zhang, Hengrui, et al.
Published: (2025)
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
A Survey of Large Language Models in Finance (FinLLMs)
by: Lee, Jean, et al.
Published: (2024)
by: Lee, Jean, et al.
Published: (2024)
In-game Toxic Language Detection: Shared Task and Attention Residuals
by: Jia, Yuanzhe, et al.
Published: (2022)
by: Jia, Yuanzhe, et al.
Published: (2022)
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
by: Han, Soyeon Caren, et al.
Published: (2024)
by: Han, Soyeon Caren, et al.
Published: (2024)
SPANER: Shared Prompt Aligner for Multimodal Semantic Representation
by: Ng, Thye Shan, et al.
Published: (2025)
by: Ng, Thye Shan, et al.
Published: (2025)
AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
Similar Items
-
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024) -
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025) -
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
by: Xiang, Biao, et al.
Published: (2026) -
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
by: Park, Jiwon, et al.
Published: (2025) -
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)