RAVEN: Multitask Retrieval Augmented Vision-Language Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rao, Varun Nagaraj, Choudhary, Siddharth, Deshpande, Aditya, Satzoda, Ravi Kumar, Appalaraju, Srikar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
von: Long, Zijun, et al.
Veröffentlicht: (2025)
von: Long, Zijun, et al.
Veröffentlicht: (2025)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Open-SAT: LLM-Guided Query Embedding Refinement for Open-Vocabulary Object Retrieval in Satellite Imagery
von: Arefeen, Md Adnan, et al.
Veröffentlicht: (2026)
von: Arefeen, Md Adnan, et al.
Veröffentlicht: (2026)
Reasoning-Augmented Representations for Multimodal Retrieval
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
von: Meng, GuangHao, et al.
Veröffentlicht: (2025)
von: Meng, GuangHao, et al.
Veröffentlicht: (2025)
TalentMine: LLM-Based Extraction and Question-Answering from Multimodal Talent Tables
von: Mannam, Varun, et al.
Veröffentlicht: (2025)
von: Mannam, Varun, et al.
Veröffentlicht: (2025)
Open Multimodal Retrieval-Augmented Factual Image Generation
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Multi-task Cross-modal Learning for Chest X-ray Image Retrieval
von: Liang, Zhaohui, et al.
Veröffentlicht: (2026)
von: Liang, Zhaohui, et al.
Veröffentlicht: (2026)
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
von: Waseda, Futa, et al.
Veröffentlicht: (2024)
von: Waseda, Futa, et al.
Veröffentlicht: (2024)
Compressible and Searchable: AI-native Multi-Modal Retrieval System with Learned Image Compression
von: Luo, Jixiang
Veröffentlicht: (2024)
von: Luo, Jixiang
Veröffentlicht: (2024)
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
von: Williams-Lekuona, Mikel, et al.
Veröffentlicht: (2025)
von: Williams-Lekuona, Mikel, et al.
Veröffentlicht: (2025)
RANa: Retrieval-Augmented Navigation
von: Monaci, Gianluca, et al.
Veröffentlicht: (2025)
von: Monaci, Gianluca, et al.
Veröffentlicht: (2025)
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation
von: Ling, Run, et al.
Veröffentlicht: (2025)
von: Ling, Run, et al.
Veröffentlicht: (2025)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
von: Askari, Arian, et al.
Veröffentlicht: (2025)
von: Askari, Arian, et al.
Veröffentlicht: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
von: Chakraborty, Debashish, et al.
Veröffentlicht: (2026)
von: Chakraborty, Debashish, et al.
Veröffentlicht: (2026)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
von: Yang, Bang, et al.
Veröffentlicht: (2024)
von: Yang, Bang, et al.
Veröffentlicht: (2024)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
A Multi-Stage Hybrid Framework for Automated Interpretation of Multi-View Engineering Drawings Using Vision Language Model
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
Identity-Decoupled Anonymization for Visual Evidence in Multi-modal Retrieval-Augmented Generation
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
Benchmark Granularity and Model Robustness for Image-Text Retrieval
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
Weakly Supervised Deep Hyperspherical Quantization for Image Retrieval
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
Retrieval-Guided Generation for Safer Histopathology Image Captioning
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
Geometric Analysis of Self-Supervised Vision Representations for Semantic Image Retrieval
von: Rodríguez-Betancourt, Esteban, et al.
Veröffentlicht: (2026)
von: Rodríguez-Betancourt, Esteban, et al.
Veröffentlicht: (2026)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2024) -
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025) -
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026) -
Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025) -
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
von: Long, Zijun, et al.
Veröffentlicht: (2025)