RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Yinglu, Lu, Zhiying, Liu, Zhihang, Sun, Yiwei, Liu, Chuanbin, Xie, Hongtao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Hallucination Mitigation Prompts Long-term Video Understanding
par: Sun, Yiwei, et autres
Publié: (2024)
par: Sun, Yiwei, et autres
Publié: (2024)
From Evaluation to Defense: Advancing Safety in Video Large Language Models
par: Sun, Yiwei, et autres
Publié: (2025)
par: Sun, Yiwei, et autres
Publié: (2025)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
par: Pu, Bowei, et autres
Publié: (2025)
par: Pu, Bowei, et autres
Publié: (2025)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
par: Cao, Linhan, et autres
Publié: (2026)
par: Cao, Linhan, et autres
Publié: (2026)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
par: Wu, Hao, et autres
Publié: (2024)
par: Wu, Hao, et autres
Publié: (2024)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
par: Sun, Yubo, et autres
Publié: (2025)
par: Sun, Yubo, et autres
Publié: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
par: Xue, Zhucun, et autres
Publié: (2025)
par: Xue, Zhucun, et autres
Publié: (2025)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
par: Zeng, Nianbo, et autres
Publié: (2025)
par: Zeng, Nianbo, et autres
Publié: (2025)
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
par: Wang, Shuai, et autres
Publié: (2025)
par: Wang, Shuai, et autres
Publié: (2025)
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
par: Xie, Peijin, et autres
Publié: (2025)
par: Xie, Peijin, et autres
Publié: (2025)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
par: Liu, Zhihang, et autres
Publié: (2025)
par: Liu, Zhihang, et autres
Publié: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
par: Tanaka, Ryota, et autres
Publié: (2025)
par: Tanaka, Ryota, et autres
Publié: (2025)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
par: Liu, Zhihang, et autres
Publié: (2025)
par: Liu, Zhihang, et autres
Publié: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
par: Wu, Yin, et autres
Publié: (2025)
par: Wu, Yin, et autres
Publié: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
par: Chen, Jian, et autres
Publié: (2025)
par: Chen, Jian, et autres
Publié: (2025)
PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
par: Gao, Yifan, et autres
Publié: (2025)
par: Gao, Yifan, et autres
Publié: (2025)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
par: Wu, Yuxuan, et autres
Publié: (2025)
par: Wu, Yuxuan, et autres
Publié: (2025)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
par: Chen, I-Hsiang, et autres
Publié: (2026)
par: Chen, I-Hsiang, et autres
Publié: (2026)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
par: Wang, Qiuchen, et autres
Publié: (2025)
par: Wang, Qiuchen, et autres
Publié: (2025)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
par: Loo, Gowen, et autres
Publié: (2025)
par: Loo, Gowen, et autres
Publié: (2025)
NICO-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Understanding the Nicotine Public Health Crisis
par: Serna-Aguilera, Manuel, et autres
Publié: (2026)
par: Serna-Aguilera, Manuel, et autres
Publié: (2026)
Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
par: Sun, Xiatao, et autres
Publié: (2025)
par: Sun, Xiatao, et autres
Publié: (2025)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
par: Qi, Jingyuan, et autres
Publié: (2025)
par: Qi, Jingyuan, et autres
Publié: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
par: Wang, Qiuchen, et autres
Publié: (2026)
par: Wang, Qiuchen, et autres
Publié: (2026)
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
par: Li, Gen, et autres
Publié: (2026)
par: Li, Gen, et autres
Publié: (2026)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
par: Zhu, Chenhui, et autres
Publié: (2025)
par: Zhu, Chenhui, et autres
Publié: (2025)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
par: Li, Peize, et autres
Publié: (2026)
par: Li, Peize, et autres
Publié: (2026)
FairRAG: Fair Human Generation via Fair Retrieval Augmentation
par: Shrestha, Robik, et autres
Publié: (2024)
par: Shrestha, Robik, et autres
Publié: (2024)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
par: Wang, Wenbin, et autres
Publié: (2025)
par: Wang, Wenbin, et autres
Publié: (2025)
Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering
par: Naeem, Awais, et autres
Publié: (2024)
par: Naeem, Awais, et autres
Publié: (2024)
ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
par: Liu, Zhihang, et autres
Publié: (2025)
par: Liu, Zhihang, et autres
Publié: (2025)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
par: Zhang, Junyuan, et autres
Publié: (2024)
par: Zhang, Junyuan, et autres
Publié: (2024)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
par: Luo, Yongdong, et autres
Publié: (2024)
par: Luo, Yongdong, et autres
Publié: (2024)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
par: Lu, Songshuo, et autres
Publié: (2024)
par: Lu, Songshuo, et autres
Publié: (2024)
TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
par: Cao, Zongsheng, et autres
Publié: (2025)
par: Cao, Zongsheng, et autres
Publié: (2025)
UniPLV: Towards Label-Efficient Open-World 3D Scene Understanding by Regional Visual Language Supervision
par: Wang, Yuru, et autres
Publié: (2024)
par: Wang, Yuru, et autres
Publié: (2024)
RegionGPT: Towards Region Understanding Vision Language Model
par: Guo, Qiushan, et autres
Publié: (2024)
par: Guo, Qiushan, et autres
Publié: (2024)
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
par: Xue, Junxiao, et autres
Publié: (2026)
par: Xue, Junxiao, et autres
Publié: (2026)
AstroRAG -- A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy
par: Wang, Zhifeng, et autres
Publié: (2026)
par: Wang, Zhifeng, et autres
Publié: (2026)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
par: Wang, Jiankang, et autres
Publié: (2025)
par: Wang, Jiankang, et autres
Publié: (2025)
Documents similaires
-
Hallucination Mitigation Prompts Long-term Video Understanding
par: Sun, Yiwei, et autres
Publié: (2024) -
From Evaluation to Defense: Advancing Safety in Video Large Language Models
par: Sun, Yiwei, et autres
Publié: (2025) -
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
par: Pu, Bowei, et autres
Publié: (2025) -
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
par: Cao, Linhan, et autres
Publié: (2026) -
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
par: Wu, Hao, et autres
Publié: (2024)