SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Yihao, Han, Soyeon Caren, Li, Zechuan, Chung, Hyunsuk |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
von: Xiang, Biao, et al.
Veröffentlicht: (2026)
von: Xiang, Biao, et al.
Veröffentlicht: (2026)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
Deep Learning based Visually Rich Document Content Understanding: A Survey
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Graph Neural Networks for Text Classification: A Survey
von: Wang, Kunze, et al.
Veröffentlicht: (2023)
von: Wang, Kunze, et al.
Veröffentlicht: (2023)
PEACH: Pretrained-embedding Explanation Across Contextual and Hierarchical Structure
von: Cao, Feiqi, et al.
Veröffentlicht: (2024)
von: Cao, Feiqi, et al.
Veröffentlicht: (2024)
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
von: Li, Yan, et al.
Veröffentlicht: (2024)
von: Li, Yan, et al.
Veröffentlicht: (2024)
Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs
von: Pratama, Dhita Putri, et al.
Veröffentlicht: (2026)
von: Pratama, Dhita Putri, et al.
Veröffentlicht: (2026)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
von: Akasaka, Muku, et al.
Veröffentlicht: (2026)
von: Akasaka, Muku, et al.
Veröffentlicht: (2026)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
von: Park, Jiwon, et al.
Veröffentlicht: (2025)
von: Park, Jiwon, et al.
Veröffentlicht: (2025)
ToolTree: Efficient LLM Agent Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
Physics-based phenomenological characterization of cross-modal bias in multimodal models
von: Kim, Hyeongmo, et al.
Veröffentlicht: (2026)
von: Kim, Hyeongmo, et al.
Veröffentlicht: (2026)
JAC-Antimicrobial Resistance
Veröffentlicht: (2020)
Veröffentlicht: (2020)
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
von: Zhao, Haoran, et al.
Veröffentlicht: (2026)
von: Zhao, Haoran, et al.
Veröffentlicht: (2026)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
von: Dai, Yue, et al.
Veröffentlicht: (2025)
von: Dai, Yue, et al.
Veröffentlicht: (2025)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
UniCast: A Unified Framework for Instance-Conditioned Multimodal Time-Series Forecasting
von: Park, Sehyuk, et al.
Veröffentlicht: (2025)
von: Park, Sehyuk, et al.
Veröffentlicht: (2025)
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU
von: Li, Yan, et al.
Veröffentlicht: (2024)
von: Li, Yan, et al.
Veröffentlicht: (2024)
FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance
von: Chung, Hyunsuk, et al.
Veröffentlicht: (2026)
von: Chung, Hyunsuk, et al.
Veröffentlicht: (2026)
SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction
von: Bradley, Ethan, et al.
Veröffentlicht: (2024)
von: Bradley, Ethan, et al.
Veröffentlicht: (2024)
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2026)
von: Ding, Yihao, et al.
Veröffentlicht: (2026)
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
von: Ng, Thye Shan, et al.
Veröffentlicht: (2024)
von: Ng, Thye Shan, et al.
Veröffentlicht: (2024)
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
A Survey of Large Language Models in Finance (FinLLMs)
von: Lee, Jean, et al.
Veröffentlicht: (2024)
von: Lee, Jean, et al.
Veröffentlicht: (2024)
In-game Toxic Language Detection: Shared Task and Attention Residuals
von: Jia, Yuanzhe, et al.
Veröffentlicht: (2022)
von: Jia, Yuanzhe, et al.
Veröffentlicht: (2022)
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
von: Han, Soyeon Caren, et al.
Veröffentlicht: (2024)
von: Han, Soyeon Caren, et al.
Veröffentlicht: (2024)
SPANER: Shared Prompt Aligner for Multimodal Semantic Representation
von: Ng, Thye Shan, et al.
Veröffentlicht: (2025)
von: Ng, Thye Shan, et al.
Veröffentlicht: (2025)
Exploring Syn-to-Real Domain Adaptation for Military Target Detection
von: Jeong, Jongoh, et al.
Veröffentlicht: (2025)
von: Jeong, Jongoh, et al.
Veröffentlicht: (2025)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
von: Mishra, Samarth, et al.
Veröffentlicht: (2023)
von: Mishra, Samarth, et al.
Veröffentlicht: (2023)
Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
von: Liang, Zhili Nicholas, et al.
Veröffentlicht: (2026)
von: Liang, Zhili Nicholas, et al.
Veröffentlicht: (2026)
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
von: Li, Haoyi, et al.
Veröffentlicht: (2025)
von: Li, Haoyi, et al.
Veröffentlicht: (2025)
MUSEKG: A Knowledge Graph Over Museum Collections
von: Li, Jinhao, et al.
Veröffentlicht: (2025)
von: Li, Jinhao, et al.
Veröffentlicht: (2025)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
von: Cabral, Rina Carines, et al.
Veröffentlicht: (2024)
von: Cabral, Rina Carines, et al.
Veröffentlicht: (2024)
EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
von: Yuan, Angela Yifei, et al.
Veröffentlicht: (2025)
von: Yuan, Angela Yifei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
von: Ding, Yihao, et al.
Veröffentlicht: (2025) -
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
von: Xiang, Biao, et al.
Veröffentlicht: (2026) -
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2025) -
Deep Learning based Visually Rich Document Content Understanding: A Survey
von: Ding, Yihao, et al.
Veröffentlicht: (2024) -
Graph Neural Networks for Text Classification: A Survey
von: Wang, Kunze, et al.
Veröffentlicht: (2023)