URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Yongxin, Wang, Jiapeng, Shan, Zeyu, Peng, Dezhi, Lin, Zening, Jin, Lianwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
C$^{3}$Bench: A Comprehensive Classical Chinese Understanding Benchmark for Large Language Models
di: Cao, Jiahuan, et al.
Pubblicazione: (2024)
di: Cao, Jiahuan, et al.
Pubblicazione: (2024)
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction
di: Lin, Zening, et al.
Pubblicazione: (2024)
di: Lin, Zening, et al.
Pubblicazione: (2024)
TongGu: Mastering Classical Chinese Understanding with Knowledge-Grounded Large Language Models
di: Cao, Jiahuan, et al.
Pubblicazione: (2024)
di: Cao, Jiahuan, et al.
Pubblicazione: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
Predicting the Original Appearance of Damaged Historical Documents
di: Yang, Zhenhua, et al.
Pubblicazione: (2024)
di: Yang, Zhenhua, et al.
Pubblicazione: (2024)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
UPOCR: Towards Unified Pixel-Level OCR Interface
di: Peng, Dezhi, et al.
Pubblicazione: (2023)
di: Peng, Dezhi, et al.
Pubblicazione: (2023)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
di: Chen, Yen-Shan, et al.
Pubblicazione: (2024)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2024)
Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
di: Zhang, Yuyi, et al.
Pubblicazione: (2025)
di: Zhang, Yuyi, et al.
Pubblicazione: (2025)
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
di: Liao, Wenhui, et al.
Pubblicazione: (2024)
di: Liao, Wenhui, et al.
Pubblicazione: (2024)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs
di: Zhang, Jintian, et al.
Pubblicazione: (2024)
di: Zhang, Jintian, et al.
Pubblicazione: (2024)
Unified Multimodal Interleaved Document Representation for Retrieval
di: Lee, Jaewoo, et al.
Pubblicazione: (2024)
di: Lee, Jaewoo, et al.
Pubblicazione: (2024)
LongFin: A Multimodal Document Understanding Model for Long Financial Domain Documents
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding
di: Li, Yuqing, et al.
Pubblicazione: (2025)
di: Li, Yuqing, et al.
Pubblicazione: (2025)
Clustering Algorithms and RAG Enhancing Semi-Supervised Text Classification with Large LLMs
di: Zhong, Shan, et al.
Pubblicazione: (2024)
di: Zhong, Shan, et al.
Pubblicazione: (2024)
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
Hierarchical Document Refinement for Long-context Retrieval-augmented Generation
di: Jin, Jiajie, et al.
Pubblicazione: (2025)
di: Jin, Jiajie, et al.
Pubblicazione: (2025)
Structured Attention Matters to Multimodal LLMs in Document Understanding
di: Liu, Chang, et al.
Pubblicazione: (2025)
di: Liu, Chang, et al.
Pubblicazione: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
di: Gao, Sensen, et al.
Pubblicazione: (2025)
di: Gao, Sensen, et al.
Pubblicazione: (2025)
MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation
di: Zhang, Yongyue, et al.
Pubblicazione: (2026)
di: Zhang, Yongyue, et al.
Pubblicazione: (2026)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
di: Deng, Chao, et al.
Pubblicazione: (2024)
di: Deng, Chao, et al.
Pubblicazione: (2024)
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
di: Wang, Chenlong, et al.
Pubblicazione: (2026)
di: Wang, Chenlong, et al.
Pubblicazione: (2026)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
di: Luo, Chuwei, et al.
Pubblicazione: (2022)
di: Luo, Chuwei, et al.
Pubblicazione: (2022)
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
di: Zhang, Qianchi, et al.
Pubblicazione: (2025)
di: Zhang, Qianchi, et al.
Pubblicazione: (2025)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
di: Wang, Yongxin, et al.
Pubblicazione: (2024)
di: Wang, Yongxin, et al.
Pubblicazione: (2024)
PersonaVLM: Long-Term Personalized Multimodal LLMs
di: Nie, Chang, et al.
Pubblicazione: (2026)
di: Nie, Chang, et al.
Pubblicazione: (2026)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
di: Jia, Junlong, et al.
Pubblicazione: (2025)
di: Jia, Junlong, et al.
Pubblicazione: (2025)
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
di: Peng, Dezhi, et al.
Pubblicazione: (2023)
di: Peng, Dezhi, et al.
Pubblicazione: (2023)
XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
di: Xiao, Yuzhuo, et al.
Pubblicazione: (2025)
di: Xiao, Yuzhuo, et al.
Pubblicazione: (2025)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
di: Ye, Xiaoju, et al.
Pubblicazione: (2025)
di: Ye, Xiaoju, et al.
Pubblicazione: (2025)
Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs
di: Yue, Chongjian, et al.
Pubblicazione: (2023)
di: Yue, Chongjian, et al.
Pubblicazione: (2023)
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
di: Tian, Yufei, et al.
Pubblicazione: (2024)
di: Tian, Yufei, et al.
Pubblicazione: (2024)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
di: Lin, Sheng-Chieh, et al.
Pubblicazione: (2024)
di: Lin, Sheng-Chieh, et al.
Pubblicazione: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
C$^{3}$Bench: A Comprehensive Classical Chinese Understanding Benchmark for Large Language Models
di: Cao, Jiahuan, et al.
Pubblicazione: (2024) -
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction
di: Lin, Zening, et al.
Pubblicazione: (2024) -
TongGu: Mastering Classical Chinese Understanding with Knowledge-Grounded Large Language Models
di: Cao, Jiahuan, et al.
Pubblicazione: (2024) -
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
di: Wang, Jiapeng, et al.
Pubblicazione: (2024) -
Predicting the Original Appearance of Damaged Historical Documents
di: Yang, Zhenhua, et al.
Pubblicazione: (2024)