Structured Attention Matters to Multimodal LLMs in Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Chang, Chen, Hongkai, Cai, Yujun, Wu, Hang, Ye, Qingwen, Yang, Ming-Hsuan, Wang, Yiwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Manager: Parallel Agent Loop for Long-form Deep Research
by: Xu, Yilong, et al.
Published: (2026)
by: Xu, Yilong, et al.
Published: (2026)
LITA: An Efficient LLM-assisted Iterative Topic Augmentation Framework
by: Chang, Chia-Hsuan, et al.
Published: (2024)
by: Chang, Chia-Hsuan, et al.
Published: (2024)
Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
by: Lan, Junwei, et al.
Published: (2025)
by: Lan, Junwei, et al.
Published: (2025)
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
by: Wu, Hui, et al.
Published: (2026)
by: Wu, Hui, et al.
Published: (2026)
Unified Multimodal Interleaved Document Representation for Retrieval
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation
by: Liu, Hung Ming
Published: (2026)
by: Liu, Hung Ming
Published: (2026)
M-$LLM^3$REC: A Motivation-Aware User-Item Interaction Framework for Enhancing Recommendation Accuracy with LLMs
by: Chen, Lining, et al.
Published: (2025)
by: Chen, Lining, et al.
Published: (2025)
Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents
by: Zhuang, Jiaxi, et al.
Published: (2025)
by: Zhuang, Jiaxi, et al.
Published: (2025)
Multimodal Quantitative Language for Generative Recommendation
by: Zhai, Jianyang, et al.
Published: (2025)
by: Zhai, Jianyang, et al.
Published: (2025)
LLM4Rec: Large Language Models for Multimodal Generative Recommendation with Causal Debiasing
by: Ma, Bo, et al.
Published: (2025)
by: Ma, Bo, et al.
Published: (2025)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
by: Hu, Caiyu, et al.
Published: (2025)
by: Hu, Caiyu, et al.
Published: (2025)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
by: Khanghah, Kiarash Naghavi, et al.
Published: (2026)
by: Khanghah, Kiarash Naghavi, et al.
Published: (2026)
Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
by: Su, Weihang, et al.
Published: (2026)
by: Su, Weihang, et al.
Published: (2026)
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
by: Wu, Kaifeng, et al.
Published: (2025)
by: Wu, Kaifeng, et al.
Published: (2025)
AttentionRetriever: Attention Layers are Secretly Long Document Retrievers
by: Fu, David Jiahao, et al.
Published: (2026)
by: Fu, David Jiahao, et al.
Published: (2026)
Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI
by: Singh, Saurabh K., et al.
Published: (2026)
by: Singh, Saurabh K., et al.
Published: (2026)
Tuning LLMs by RAG Principles: Towards LLM-native Memory
by: Wei, Jiale, et al.
Published: (2025)
by: Wei, Jiale, et al.
Published: (2025)
Retrieval-Augmented LLMs for Evidence Localization in Clinical Trial Recruitment from Longitudinal EHR Narratives
by: Chen, Ziyi, et al.
Published: (2026)
by: Chen, Ziyi, et al.
Published: (2026)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
DeepRead: Document Structure-Aware Reasoning to Enhance Agentic Search
by: Li, Zhanli, et al.
Published: (2026)
by: Li, Zhanli, et al.
Published: (2026)
Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
Quantifying Document Impact in RAG-LLMs
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
by: Dai, Ji, et al.
Published: (2026)
by: Dai, Ji, et al.
Published: (2026)
LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models
by: Yang, Hang, et al.
Published: (2024)
by: Yang, Hang, et al.
Published: (2024)
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
by: Wu, Eric, et al.
Published: (2024)
by: Wu, Eric, et al.
Published: (2024)
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
by: Tavakoli, Mohammad, et al.
Published: (2025)
by: Tavakoli, Mohammad, et al.
Published: (2025)
Model-Document Protocol for AI Search
by: Qian, Hongjin, et al.
Published: (2025)
by: Qian, Hongjin, et al.
Published: (2025)
TableRAG: Million-Token Table Understanding with Language Models
by: Chen, Si-An, et al.
Published: (2024)
by: Chen, Si-An, et al.
Published: (2024)
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
by: Chen, Huiyao, et al.
Published: (2025)
by: Chen, Huiyao, et al.
Published: (2025)
CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs
by: Keerthana, Garapati, et al.
Published: (2025)
by: Keerthana, Garapati, et al.
Published: (2025)
Optimizing Multi-Hop Document Retrieval Through Intermediate Representations
by: Lin, Jiaen, et al.
Published: (2025)
by: Lin, Jiaen, et al.
Published: (2025)
Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs
by: Duwal, Sharad, et al.
Published: (2025)
by: Duwal, Sharad, et al.
Published: (2025)
Do LLMs Understand Collaborative Signals? Diagnosis and Repair
by: Pouryousef, Shahrooz, et al.
Published: (2025)
by: Pouryousef, Shahrooz, et al.
Published: (2025)
Beyond path selection: Better LLMs for Scientific Information Extraction with MimicSFT and Relevance and Rule-induced(R$^2$)GRPO
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
Similar Items
-
Self-Manager: Parallel Agent Loop for Long-form Deep Research
by: Xu, Yilong, et al.
Published: (2026) -
LITA: An Efficient LLM-assisted Iterative Topic Augmentation Framework
by: Chang, Chia-Hsuan, et al.
Published: (2024) -
Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
by: Lan, Junwei, et al.
Published: (2025) -
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
by: Wu, Hui, et al.
Published: (2026) -
Unified Multimodal Interleaved Document Representation for Retrieval
by: Lee, Jaewoo, et al.
Published: (2024)