M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Haolong, Tan, Kaijun, Shen, Yeqing, Huang, Xin, Ge, Zheng, Zhang, Xiangyu, Li, Si, Jiang, Daxin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
by: Yan, Haolong, et al.
Published: (2025)
by: Yan, Haolong, et al.
Published: (2025)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
by: Fan, Cunxin, et al.
Published: (2025)
by: Fan, Cunxin, et al.
Published: (2025)
Towards Text-Image Interleaved Retrieval
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Interleaving Reasoning for Better Text-to-Image Generation
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
by: Zhou, Chenyu, et al.
Published: (2024)
by: Zhou, Chenyu, et al.
Published: (2024)
DocTer: Documentation Guided Fuzzing for Testing Deep Learning API Functions
by: Xie, Danning, et al.
Published: (2021)
by: Xie, Danning, et al.
Published: (2021)
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)
by: Purkayastha, Kunal, et al.
Published: (2026)
Holistic Evaluation for Interleaved Text-and-Image Generation
by: Liu, Minqian, et al.
Published: (2024)
by: Liu, Minqian, et al.
Published: (2024)
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
Comprehending Columbine
by: Larkin, Ralph
Published: (2017)
by: Larkin, Ralph
Published: (2017)
Teach Multimodal LLMs to Comprehend Electrocardiographic Images
by: Liu, Ruoqi, et al.
Published: (2024)
by: Liu, Ruoqi, et al.
Published: (2024)
SumRank: Aligning Summarization Models for Long-Document Listwise Reranking
by: Feng, Jincheng, et al.
Published: (2026)
by: Feng, Jincheng, et al.
Published: (2026)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs
by: Ding, Zhikai, et al.
Published: (2025)
by: Ding, Zhikai, et al.
Published: (2025)
Do Multi-Document Summarization Models Synthesize?
by: DeYoung, Jay, et al.
Published: (2023)
by: DeYoung, Jay, et al.
Published: (2023)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
by: Kou, Siqi, et al.
Published: (2024)
by: Kou, Siqi, et al.
Published: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024)
by: Tian, Changyao, et al.
Published: (2024)
AgenticSum: An Agentic Inference-Time Framework for Faithful Clinical Text Summarization
by: Piya, Fahmida Liza, et al.
Published: (2026)
by: Piya, Fahmida Liza, et al.
Published: (2026)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
DocReward: A Document Reward Model for Structuring and Stylizing
by: Liu, Junpeng, et al.
Published: (2025)
by: Liu, Junpeng, et al.
Published: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
by: Xue, Zhenliang, et al.
Published: (2025)
by: Xue, Zhenliang, et al.
Published: (2025)
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
by: Zheng, Ge, et al.
Published: (2025)
by: Zheng, Ge, et al.
Published: (2025)
MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking
by: Chen, Ting-Chih, et al.
Published: (2024)
by: Chen, Ting-Chih, et al.
Published: (2024)
StrucSum: Graph-Structured Reasoning for Long Document Extractive Summarization with LLMs
by: Yuan, Haohan, et al.
Published: (2025)
by: Yuan, Haohan, et al.
Published: (2025)
Comprehending and Confronting Antisemitism
Published: (2020)
Published: (2020)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
by: Cao, Junjie, et al.
Published: (2025)
by: Cao, Junjie, et al.
Published: (2025)
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
by: Shu, Bao, et al.
Published: (2025)
by: Shu, Bao, et al.
Published: (2025)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
by: Shen, Meng, et al.
Published: (2026)
by: Shen, Meng, et al.
Published: (2026)
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
by: Zheng, Haolong, et al.
Published: (2025)
by: Zheng, Haolong, et al.
Published: (2025)
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
by: Zou, Zhentao, et al.
Published: (2025)
by: Zou, Zhentao, et al.
Published: (2025)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
by: Cui, Wenqian, et al.
Published: (2026)
by: Cui, Wenqian, et al.
Published: (2026)
A Method of Extractive Text Summarization Using Document Semantic Graph With Node Ranking
by: Zhenhao Li, et al.
Published: (2025)
by: Zhenhao Li, et al.
Published: (2025)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
Positional Cognitive Specialization: Where Do LLMs Learn To Comprehend and Speak Your Language?
by: Salim, Luis Frentzen, et al.
Published: (2026)
by: Salim, Luis Frentzen, et al.
Published: (2026)
VisDocSketcher: Towards Scalable Visual Documentation with Agentic Systems
by: Gomes, Luís F., et al.
Published: (2025)
by: Gomes, Luís F., et al.
Published: (2025)
Comprehending the experience of being anxious
by: Alberto Mario de Castro
Published: (2004)
by: Alberto Mario de Castro
Published: (2004)
DomainSum: A Hierarchical Benchmark for Fine-Grained Domain Shift in Abstractive Text Summarization
by: Yuan, Haohan, et al.
Published: (2024)
by: Yuan, Haohan, et al.
Published: (2024)
Similar Items
-
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
by: Chau, Phan Phuong Mai, et al.
Published: (2024) -
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
by: Yan, Haolong, et al.
Published: (2025) -
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
by: Fan, Cunxin, et al.
Published: (2025) -
Towards Text-Image Interleaved Retrieval
by: Zhang, Xin, et al.
Published: (2025) -
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
by: Chen, Dongping, et al.
Published: (2024)