Attacks on multimodal models
Fuente:
arXiv
Saved in:
| Main Authors: | Iablochnikov, Viacheslav, Rogachev, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Generate Rigid Body Interactions with Video Diffusion Models
by: Romero, David, et al.
Published: (2025)
by: Romero, David, et al.
Published: (2025)
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
by: Chen, Yanyuan, et al.
Published: (2025)
by: Chen, Yanyuan, et al.
Published: (2025)
SalsaAgent: A multimodal embodied language model for interactive dance generation
by: Yazdian, Payam Jome, et al.
Published: (2026)
by: Yazdian, Payam Jome, et al.
Published: (2026)
Monitoring Horses in Stalls: From Object to Event Detection
by: Galimzianov, Dmitrii, et al.
Published: (2025)
by: Galimzianov, Dmitrii, et al.
Published: (2025)
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
by: Chae, Joongwon, et al.
Published: (2024)
by: Chae, Joongwon, et al.
Published: (2024)
In-context learning enables multimodal large language models to classify cancer pathology images
by: Ferber, Dyke, et al.
Published: (2024)
by: Ferber, Dyke, et al.
Published: (2024)
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
by: Pang, Yik Lung, et al.
Published: (2026)
by: Pang, Yik Lung, et al.
Published: (2026)
SmolVLM: Redefining small and efficient multimodal models
by: Marafioti, Andrés, et al.
Published: (2025)
by: Marafioti, Andrés, et al.
Published: (2025)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
by: Wu, Yiqi, et al.
Published: (2024)
by: Wu, Yiqi, et al.
Published: (2024)
Text-to-Vector Conversion for Residential Plan Design
by: Bazhenov, Egor, et al.
Published: (2026)
by: Bazhenov, Egor, et al.
Published: (2026)
Transformation trees -- documentation of multimodal image registration
by: Tomaka, Agnieszka Anna, et al.
Published: (2025)
by: Tomaka, Agnieszka Anna, et al.
Published: (2025)
Visual Language Models as Zero-Shot Deepfake Detectors
by: Pirogov, Viacheslav
Published: (2025)
by: Pirogov, Viacheslav
Published: (2025)
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Animalbooth: multimodal feature enhancement for animal subject personalization
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
Densification and forecasting of Sentinel-2 time series from multimodal SAR and Optical satellite data using deep generative models
by: Defonte, Véronique, et al.
Published: (2026)
by: Defonte, Véronique, et al.
Published: (2026)
New multimodal similarity measure for image registration via modeling local functional dependence with linear combination of learned basis functions
by: Honkamaa, Joel, et al.
Published: (2025)
by: Honkamaa, Joel, et al.
Published: (2025)
Evaluating point-light biological motion in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
Bias-constrained multimodal intelligence for equitable and reliable clinical AI
by: Li, Cheng, et al.
Published: (2026)
by: Li, Cheng, et al.
Published: (2026)
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
by: Granite Vision Team, et al.
Published: (2025)
by: Granite Vision Team, et al.
Published: (2025)
AI-powered multimodal modeling of personalized hemodynamics in aortic stenosis
by: Ozturk, Caglar, et al.
Published: (2024)
by: Ozturk, Caglar, et al.
Published: (2024)
Headset: Human emotion awareness under partial occlusions multimodal dataset
by: Lohesara, Fatemeh Ghorbani, et al.
Published: (2024)
by: Lohesara, Fatemeh Ghorbani, et al.
Published: (2024)
JUMP: A joint multimodal registration pipeline for neuroimaging with minimal preprocessing
by: Casamitjana, Adria, et al.
Published: (2024)
by: Casamitjana, Adria, et al.
Published: (2024)
From Image to Video, what do we need in multimodal LLMs?
by: Huang, Suyuan, et al.
Published: (2024)
by: Huang, Suyuan, et al.
Published: (2024)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
by: Janizek, Joseph D., et al.
Published: (2026)
by: Janizek, Joseph D., et al.
Published: (2026)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation
by: Muhovič, Jon, et al.
Published: (2025)
by: Muhovič, Jon, et al.
Published: (2025)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM
by: Yoshida, Takero, et al.
Published: (2024)
by: Yoshida, Takero, et al.
Published: (2024)
MObyGaze: a film dataset of multimodal objectification densely annotated by experts
by: Tores, Julie, et al.
Published: (2025)
by: Tores, Julie, et al.
Published: (2025)
LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
by: Zhang, Sheng, et al.
Published: (2023)
by: Zhang, Sheng, et al.
Published: (2023)
What do vision-language models see in the context? Investigating multimodal in-context learning
by: Santos, Gabriel O. dos, et al.
Published: (2025)
by: Santos, Gabriel O. dos, et al.
Published: (2025)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
by: Meng, Debin, et al.
Published: (2024)
by: Meng, Debin, et al.
Published: (2024)
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Similar Items
-
Learning to Generate Rigid Body Interactions with Video Diffusion Models
by: Romero, David, et al.
Published: (2025) -
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025) -
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
by: Chen, Yanyuan, et al.
Published: (2025) -
SalsaAgent: A multimodal embodied language model for interactive dance generation
by: Yazdian, Payam Jome, et al.
Published: (2026) -
Monitoring Horses in Stalls: From Object to Event Detection
by: Galimzianov, Dmitrii, et al.
Published: (2025)