VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guan, Guowei, Hao, Yurong, Zhang, Jiaming, Wu, Tiantong, Zhang, Fuyao, Chen, Tianxiang, Huang, Longtao, Leung, Cyril, Lim, Wei Yang Bryan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911427259793408
author Guan, Guowei
Hao, Yurong
Zhang, Jiaming
Wu, Tiantong
Zhang, Fuyao
Chen, Tianxiang
Huang, Longtao
Leung, Cyril
Lim, Wei Yang Bryan
author_facet Guan, Guowei
Hao, Yurong
Zhang, Jiaming
Wu, Tiantong
Zhang, Fuyao
Chen, Tianxiang
Huang, Longtao
Leung, Cyril
Lim, Wei Yang Bryan
contents Multimodal large language models (MLLMs) are pushing recommender systems (RecSys) toward content-grounded retrieval and ranking via cross-modal fusion. We find that while cross-modal consensus often mitigates conventional poisoning that manipulates interaction logs or perturbs a single modality, it also introduces a new attack surface where synchronised multimodal poisoning can reliably steer fused representations along stable semantic directions during fine-tuning. To characterise this threat, we formalise cross-modal interactive poisoning and propose VENOMREC, which performs Exposure Alignment to identify high-exposure regions in the joint embedding space and Cross-modal Interactive Perturbation to craft attention-guided coupled token-patch edits. Experiments on three real-world multimodal datasets demonstrate that VENOMREC consistently outperforms strong baselines, achieving 0.73 mean ER@20 and improving over the strongest baseline by +0.52 absolute ER points on average, while maintaining comparable recommendation utility.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06409
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems
Guan, Guowei
Hao, Yurong
Zhang, Jiaming
Wu, Tiantong
Zhang, Fuyao
Chen, Tianxiang
Huang, Longtao
Leung, Cyril
Lim, Wei Yang Bryan
Cryptography and Security
Multimodal large language models (MLLMs) are pushing recommender systems (RecSys) toward content-grounded retrieval and ranking via cross-modal fusion. We find that while cross-modal consensus often mitigates conventional poisoning that manipulates interaction logs or perturbs a single modality, it also introduces a new attack surface where synchronised multimodal poisoning can reliably steer fused representations along stable semantic directions during fine-tuning. To characterise this threat, we formalise cross-modal interactive poisoning and propose VENOMREC, which performs Exposure Alignment to identify high-exposure regions in the joint embedding space and Cross-modal Interactive Perturbation to craft attention-guided coupled token-patch edits. Experiments on three real-world multimodal datasets demonstrate that VENOMREC consistently outperforms strong baselines, achieving 0.73 mean ER@20 and improving over the strongest baseline by +0.52 absolute ER points on average, while maintaining comparable recommendation utility.
title VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems
topic Cryptography and Security
url https://arxiv.org/abs/2602.06409