How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Qian, Yusu, Zhang, Haotian, Yang, Yinfei, Gan, Zhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
di: Qian, Yusu, et al.
Pubblicazione: (2024)
di: Qian, Yusu, et al.
Pubblicazione: (2024)
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
di: Amirloo, Elmira, et al.
Pubblicazione: (2024)
di: Amirloo, Elmira, et al.
Pubblicazione: (2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
di: You, Keen, et al.
Pubblicazione: (2024)
di: You, Keen, et al.
Pubblicazione: (2024)
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
di: Qian, Yusu, et al.
Pubblicazione: (2025)
di: Qian, Yusu, et al.
Pubblicazione: (2025)
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
di: Qian, Yusu, et al.
Pubblicazione: (2025)
di: Qian, Yusu, et al.
Pubblicazione: (2025)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
di: Fu, Tsu-Jui, et al.
Pubblicazione: (2025)
di: Fu, Tsu-Jui, et al.
Pubblicazione: (2025)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
di: Feng, Di, et al.
Pubblicazione: (2025)
di: Feng, Di, et al.
Pubblicazione: (2025)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
di: Qian, Yusu, et al.
Pubblicazione: (2025)
di: Qian, Yusu, et al.
Pubblicazione: (2025)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
di: Zhao, Xiangyu, et al.
Pubblicazione: (2023)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2023)
FoolSDEdit: Deceptively Steering Your Edits Towards Targeted Attribute-aware Distribution
di: Zhou, Qi, et al.
Pubblicazione: (2024)
di: Zhou, Qi, et al.
Pubblicazione: (2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
di: Li, Zhangheng, et al.
Pubblicazione: (2024)
di: Li, Zhangheng, et al.
Pubblicazione: (2024)
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
di: Zhang, Yuechen, et al.
Pubblicazione: (2023)
di: Zhang, Yuechen, et al.
Pubblicazione: (2023)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
di: Zhang, Haotian, et al.
Pubblicazione: (2024)
di: Zhang, Haotian, et al.
Pubblicazione: (2024)
Dual-branch Prompting for Multimodal Machine Translation
di: Wang, Jie, et al.
Pubblicazione: (2025)
di: Wang, Jie, et al.
Pubblicazione: (2025)
ESG Accountability Made Easy: DocQA at Your Service
di: Mishra, Lokesh, et al.
Pubblicazione: (2023)
di: Mishra, Lokesh, et al.
Pubblicazione: (2023)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
di: Daxberger, Erik, et al.
Pubblicazione: (2025)
di: Daxberger, Erik, et al.
Pubblicazione: (2025)
How to Merge Your Multimodal Models Over Time?
di: Dziadzio, Sebastian, et al.
Pubblicazione: (2024)
di: Dziadzio, Sebastian, et al.
Pubblicazione: (2024)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
di: Kang, Caixin, et al.
Pubblicazione: (2025)
di: Kang, Caixin, et al.
Pubblicazione: (2025)
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
di: Diesendruck, Maurice, et al.
Pubblicazione: (2024)
di: Diesendruck, Maurice, et al.
Pubblicazione: (2024)
MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment
di: Jia, Fankai, et al.
Pubblicazione: (2025)
di: Jia, Fankai, et al.
Pubblicazione: (2025)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
di: Fu, Tsu-Jui, et al.
Pubblicazione: (2023)
di: Fu, Tsu-Jui, et al.
Pubblicazione: (2023)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
di: Yang, Yuhao, et al.
Pubblicazione: (2025)
di: Yang, Yuhao, et al.
Pubblicazione: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
di: Guo, Zichun, et al.
Pubblicazione: (2026)
di: Guo, Zichun, et al.
Pubblicazione: (2026)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
di: Xie, Yuxuan, et al.
Pubblicazione: (2024)
di: Xie, Yuxuan, et al.
Pubblicazione: (2024)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
di: Guo, Zirun, et al.
Pubblicazione: (2024)
di: Guo, Zirun, et al.
Pubblicazione: (2024)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
di: McKinzie, Brandon, et al.
Pubblicazione: (2024)
di: McKinzie, Brandon, et al.
Pubblicazione: (2024)
PersonaVLM: Long-Term Personalized Multimodal LLMs
di: Nie, Chang, et al.
Pubblicazione: (2026)
di: Nie, Chang, et al.
Pubblicazione: (2026)
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
di: Li, Yanghao, et al.
Pubblicazione: (2025)
di: Li, Yanghao, et al.
Pubblicazione: (2025)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
di: Wang, Zirui, et al.
Pubblicazione: (2024)
di: Wang, Zirui, et al.
Pubblicazione: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
di: Ji, Yatai, et al.
Pubblicazione: (2024)
di: Ji, Yatai, et al.
Pubblicazione: (2024)
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
di: Narayan, Kartik, et al.
Pubblicazione: (2025)
di: Narayan, Kartik, et al.
Pubblicazione: (2025)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
di: Yang, Zhen, et al.
Pubblicazione: (2025)
di: Yang, Zhen, et al.
Pubblicazione: (2025)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
di: Wu, Wentao, et al.
Pubblicazione: (2023)
di: Wu, Wentao, et al.
Pubblicazione: (2023)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
di: Song, Tingyu, et al.
Pubblicazione: (2025)
di: Song, Tingyu, et al.
Pubblicazione: (2025)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
di: Ye, Hanrong, et al.
Pubblicazione: (2024)
di: Ye, Hanrong, et al.
Pubblicazione: (2024)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
di: Naseh, Ali, et al.
Pubblicazione: (2024)
di: Naseh, Ali, et al.
Pubblicazione: (2024)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
di: Mi, Hongze, et al.
Pubblicazione: (2024)
di: Mi, Hongze, et al.
Pubblicazione: (2024)
Voting-based Multimodal Automatic Deception Detection
di: Touma, Lana, et al.
Pubblicazione: (2023)
di: Touma, Lana, et al.
Pubblicazione: (2023)
Documenti analoghi
-
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
di: Qian, Yusu, et al.
Pubblicazione: (2024) -
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
di: Amirloo, Elmira, et al.
Pubblicazione: (2024) -
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
di: You, Keen, et al.
Pubblicazione: (2024) -
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
di: Qian, Yusu, et al.
Pubblicazione: (2025) -
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
di: Qian, Yusu, et al.
Pubblicazione: (2025)