DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chau, Phan Phuong Mai, Bakkali, Souhail, Doucet, Antoine |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
von: Hamed, Omar, et al.
Veröffentlicht: (2024)
von: Hamed, Omar, et al.
Veröffentlicht: (2024)
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
von: Mudet, Anthony, et al.
Veröffentlicht: (2025)
von: Mudet, Anthony, et al.
Veröffentlicht: (2025)
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023)
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
IDTrust: Deep Identity Document Quality Detection with Bandpass Filtering
von: Al-Ghadi, Musab, et al.
Veröffentlicht: (2024)
von: Al-Ghadi, Musab, et al.
Veröffentlicht: (2024)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
DocReward: A Document Reward Model for Structuring and Stylizing
von: Liu, Junpeng, et al.
Veröffentlicht: (2025)
von: Liu, Junpeng, et al.
Veröffentlicht: (2025)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025)
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025)
KhmerST: A Low-Resource Khmer Scene Text Detection and Recognition Benchmark
von: Nom, Vannkinh, et al.
Veröffentlicht: (2024)
von: Nom, Vannkinh, et al.
Veröffentlicht: (2024)
Evaluating the Impact of Khmer Font Types on Text Recognition
von: Nom, Vannkinh, et al.
Veröffentlicht: (2025)
von: Nom, Vannkinh, et al.
Veröffentlicht: (2025)
VLP: A Survey on Vision-Language Pre-training
von: Chen, Feilong, et al.
Veröffentlicht: (2022)
von: Chen, Feilong, et al.
Veröffentlicht: (2022)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
Anatomical Structure-Guided Medical Vision-Language Pre-training
von: Li, Qingqiu, et al.
Veröffentlicht: (2024)
von: Li, Qingqiu, et al.
Veröffentlicht: (2024)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
von: Ni, Feng, et al.
Veröffentlicht: (2025)
von: Ni, Feng, et al.
Veröffentlicht: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
von: Lokesh, K, et al.
Veröffentlicht: (2026)
von: Lokesh, K, et al.
Veröffentlicht: (2026)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2026)
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2026)
Sigma: Semantically Informative Pre-training for Skeleton-based Sign Language Understanding
von: Pu, Muxin, et al.
Veröffentlicht: (2025)
von: Pu, Muxin, et al.
Veröffentlicht: (2025)
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
von: Mahanta, Cristina, et al.
Veröffentlicht: (2025)
von: Mahanta, Cristina, et al.
Veröffentlicht: (2025)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
TADACap: Time-series Adaptive Domain-Aware Captioning
von: Fons, Elizabeth, et al.
Veröffentlicht: (2025)
von: Fons, Elizabeth, et al.
Veröffentlicht: (2025)
Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification
von: Ye, Yaoqin, et al.
Veröffentlicht: (2024)
von: Ye, Yaoqin, et al.
Veröffentlicht: (2024)
ANNA: Abstractive Text-to-Image Synthesis with Filtered News Captions
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2023)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2023)
An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
von: Zhong, Xinliu, et al.
Veröffentlicht: (2025)
von: Zhong, Xinliu, et al.
Veröffentlicht: (2025)
Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization
von: Islam, Md Moinul, et al.
Veröffentlicht: (2025)
von: Islam, Md Moinul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
von: Yan, Haolong, et al.
Veröffentlicht: (2025) -
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
von: Hamed, Omar, et al.
Veröffentlicht: (2024) -
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
von: Mudet, Anthony, et al.
Veröffentlicht: (2025) -
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023) -
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)