Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Yin, Long, Quanyu, Li, Jing, Yu, Jianfei, Wang, Wenya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
di: Wu, Di, et al.
Pubblicazione: (2025)
di: Wu, Di, et al.
Pubblicazione: (2025)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
di: Sun, Yubo, et al.
Pubblicazione: (2025)
di: Sun, Yubo, et al.
Pubblicazione: (2025)
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation
di: Zhao, Ran, et al.
Pubblicazione: (2026)
di: Zhao, Ran, et al.
Pubblicazione: (2026)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
di: Loo, Gowen, et al.
Pubblicazione: (2025)
di: Loo, Gowen, et al.
Pubblicazione: (2025)
Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
di: Shi, Yumeng, et al.
Pubblicazione: (2025)
di: Shi, Yumeng, et al.
Pubblicazione: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
di: Li, Jiaang, et al.
Pubblicazione: (2025)
di: Li, Jiaang, et al.
Pubblicazione: (2025)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
di: Gao, Xin, et al.
Pubblicazione: (2026)
di: Gao, Xin, et al.
Pubblicazione: (2026)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
di: Tanaka, Ryota, et al.
Pubblicazione: (2025)
di: Tanaka, Ryota, et al.
Pubblicazione: (2025)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
di: Ma, Ziyu, et al.
Pubblicazione: (2024)
di: Ma, Ziyu, et al.
Pubblicazione: (2024)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
di: Huang, Ziwei, et al.
Pubblicazione: (2024)
di: Huang, Ziwei, et al.
Pubblicazione: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
di: Zhu, Mengdan, et al.
Pubblicazione: (2025)
di: Zhu, Mengdan, et al.
Pubblicazione: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
Causality Matters: How Temporal Information Emerges in Video Language Models
di: Shi, Yumeng, et al.
Pubblicazione: (2025)
di: Shi, Yumeng, et al.
Pubblicazione: (2025)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
di: Black, Alexander, et al.
Pubblicazione: (2024)
di: Black, Alexander, et al.
Pubblicazione: (2024)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
di: Song, Steven, et al.
Pubblicazione: (2024)
di: Song, Steven, et al.
Pubblicazione: (2024)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
di: Cao, Linhan, et al.
Pubblicazione: (2026)
di: Cao, Linhan, et al.
Pubblicazione: (2026)
KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
di: Hsiao, Chi-Hsiang, et al.
Pubblicazione: (2025)
di: Hsiao, Chi-Hsiang, et al.
Pubblicazione: (2025)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
di: Wang, Junling, et al.
Pubblicazione: (2026)
di: Wang, Junling, et al.
Pubblicazione: (2026)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
di: Li, Yinglu, et al.
Pubblicazione: (2025)
di: Li, Yinglu, et al.
Pubblicazione: (2025)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
di: Kabir, Imran, et al.
Pubblicazione: (2025)
di: Kabir, Imran, et al.
Pubblicazione: (2025)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
di: Chen, I-Hsiang, et al.
Pubblicazione: (2026)
di: Chen, I-Hsiang, et al.
Pubblicazione: (2026)
KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains
di: Goswami, Parthaw, et al.
Pubblicazione: (2026)
di: Goswami, Parthaw, et al.
Pubblicazione: (2026)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2024)
di: Li, Wenyan, et al.
Pubblicazione: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
di: Song, Wei, et al.
Pubblicazione: (2025)
di: Song, Wei, et al.
Pubblicazione: (2025)
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
di: Long, Xinwei, et al.
Pubblicazione: (2025)
di: Long, Xinwei, et al.
Pubblicazione: (2025)
KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
di: Huang, Hsin-Ping, et al.
Pubblicazione: (2024)
di: Huang, Hsin-Ping, et al.
Pubblicazione: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
di: Alper, Morris, et al.
Pubblicazione: (2024)
di: Alper, Morris, et al.
Pubblicazione: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
di: Zhao, Yi, et al.
Pubblicazione: (2026)
di: Zhao, Yi, et al.
Pubblicazione: (2026)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
di: Fonseca, Rui, et al.
Pubblicazione: (2025)
di: Fonseca, Rui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
di: Wu, Di, et al.
Pubblicazione: (2025) -
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025) -
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025) -
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
di: Sun, Yubo, et al.
Pubblicazione: (2025) -
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation
di: Zhao, Ran, et al.
Pubblicazione: (2026)