Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiankun, Zeng, Shenglai, Ren, Jie, Zheng, Tianqi, Liu, Hui, Tang, Xianfeng, Chang, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912383053594624
author Zhang, Jiankun
Zeng, Shenglai
Ren, Jie
Zheng, Tianqi
Liu, Hui
Tang, Xianfeng
Liu, Hui
Chang, Yi
author_facet Zhang, Jiankun
Zeng, Shenglai
Ren, Jie
Zheng, Tianqi
Liu, Hui
Tang, Xianfeng
Liu, Hui
Chang, Yi
contents Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities. While text-based RAG privacy risks have been studied, multimodal data presents unique challenges. We provide the first systematic analysis of MRAG privacy vulnerabilities across vision-language and speech-language modalities. Using a novel compositional structured prompt attack in a black-box setting, we demonstrate how attackers can extract private information by manipulating queries. Our experiments reveal that LMMs can both directly generate outputs resembling retrieved content and produce descriptions that indirectly expose sensitive information, highlighting the urgent need for robust privacy-preserving MRAG techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13957
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
Zhang, Jiankun
Zeng, Shenglai
Ren, Jie
Zheng, Tianqi
Liu, Hui
Tang, Xianfeng
Liu, Hui
Chang, Yi
Cryptography and Security
Computation and Language
Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities. While text-based RAG privacy risks have been studied, multimodal data presents unique challenges. We provide the first systematic analysis of MRAG privacy vulnerabilities across vision-language and speech-language modalities. Using a novel compositional structured prompt attack in a black-box setting, we demonstrate how attackers can extract private information by manipulating queries. Our experiments reveal that LMMs can both directly generate outputs resembling retrieved content and produce descriptions that indirectly expose sensitive information, highlighting the urgent need for robust privacy-preserving MRAG techniques.
title Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2505.13957