Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Jiachen, Wang, Changsheng, Wang, Jiongxiao, Zhang, Yiwei, Xiao, Chaowei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis
di: Umeike, Robinson, et al.
Pubblicazione: (2025)
di: Umeike, Robinson, et al.
Pubblicazione: (2025)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
di: Hou, Zhiyi, et al.
Pubblicazione: (2025)
di: Hou, Zhiyi, et al.
Pubblicazione: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
di: Rudman, William, et al.
Pubblicazione: (2026)
di: Rudman, William, et al.
Pubblicazione: (2026)
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
di: Wang, Junxin, et al.
Pubblicazione: (2026)
di: Wang, Junxin, et al.
Pubblicazione: (2026)
On the Limitations of Vision-Language Models in Understanding Image Transforms
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
di: Li, Jianing, et al.
Pubblicazione: (2024)
di: Li, Jianing, et al.
Pubblicazione: (2024)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
di: Ji, Binbin, et al.
Pubblicazione: (2025)
di: Ji, Binbin, et al.
Pubblicazione: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
di: Oliveira, Daniel, et al.
Pubblicazione: (2026)
di: Oliveira, Daniel, et al.
Pubblicazione: (2026)
Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
di: Yasuno, Takato
Pubblicazione: (2026)
di: Yasuno, Takato
Pubblicazione: (2026)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
di: Sutton, Matthew, et al.
Pubblicazione: (2026)
di: Sutton, Matthew, et al.
Pubblicazione: (2026)
Unlocking UML Class Diagram Understanding in Vision Language Models
di: Naboichenko, Artem, et al.
Pubblicazione: (2026)
di: Naboichenko, Artem, et al.
Pubblicazione: (2026)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
di: Teo, Charlton
Pubblicazione: (2025)
di: Teo, Charlton
Pubblicazione: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
From Rule-Based Models to Deep Learning Transformers Architectures for Natural Language Processing and Sign Language Translation Systems: Survey, Taxonomy and Performance Evaluation
di: Shahin, Nada, et al.
Pubblicazione: (2024)
di: Shahin, Nada, et al.
Pubblicazione: (2024)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
di: Feng, Yichen, et al.
Pubblicazione: (2026)
di: Feng, Yichen, et al.
Pubblicazione: (2026)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
di: Altamirano-Gomez, Gerardo, et al.
Pubblicazione: (2023)
di: Altamirano-Gomez, Gerardo, et al.
Pubblicazione: (2023)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
di: Slyman, Eric, et al.
Pubblicazione: (2024)
di: Slyman, Eric, et al.
Pubblicazione: (2024)
VLM4Rec: Multimodal Semantic Representation for Recommendation with Large Vision-Language Models
di: Valencia, Ty, et al.
Pubblicazione: (2026)
di: Valencia, Ty, et al.
Pubblicazione: (2026)
myMNIST: Benchmark of PETNN, KAN, and Classical Deep Learning Models for Burmese Handwritten Digit Recognition
di: Thu, Ye Kyaw, et al.
Pubblicazione: (2026)
di: Thu, Ye Kyaw, et al.
Pubblicazione: (2026)
Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives
di: Sirotkina, Elena
Pubblicazione: (2026)
di: Sirotkina, Elena
Pubblicazione: (2026)
Yanyun-3: Enabling Cross-Platform Strategy Game Operation with Vision-Language Models
di: Wang, Guoyan, et al.
Pubblicazione: (2025)
di: Wang, Guoyan, et al.
Pubblicazione: (2025)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning
di: Vlachos, Angelos, et al.
Pubblicazione: (2025)
di: Vlachos, Angelos, et al.
Pubblicazione: (2025)
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
di: Skripkin, Matvey, et al.
Pubblicazione: (2025)
di: Skripkin, Matvey, et al.
Pubblicazione: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
GAEA: A Geolocation Aware Conversational Assistant
di: Campos, Ron, et al.
Pubblicazione: (2025)
di: Campos, Ron, et al.
Pubblicazione: (2025)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
di: Hacheme, Gilles Quentin, et al.
Pubblicazione: (2025)
di: Hacheme, Gilles Quentin, et al.
Pubblicazione: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
di: Lyu, Wenbo, et al.
Pubblicazione: (2025)
di: Lyu, Wenbo, et al.
Pubblicazione: (2025)
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding
di: Wang, Chiyue, et al.
Pubblicazione: (2026)
di: Wang, Chiyue, et al.
Pubblicazione: (2026)
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
di: Zeng, Fanwei, et al.
Pubblicazione: (2026)
di: Zeng, Fanwei, et al.
Pubblicazione: (2026)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2022)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2022)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2024)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2024)
Talking Tennis: Language Feedback from 3D Biomechanical Action Recognition
di: Dashore, Arushi, et al.
Pubblicazione: (2025)
di: Dashore, Arushi, et al.
Pubblicazione: (2025)
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
di: Yang, Qi, et al.
Pubblicazione: (2026)
di: Yang, Qi, et al.
Pubblicazione: (2026)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
Devanagari Handwritten Character Recognition using Convolutional Neural Network
di: Mehta, Diksha, et al.
Pubblicazione: (2025)
di: Mehta, Diksha, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis
di: Umeike, Robinson, et al.
Pubblicazione: (2025) -
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
di: Hou, Zhiyi, et al.
Pubblicazione: (2025) -
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
di: Rudman, William, et al.
Pubblicazione: (2026) -
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
di: Wang, Junxin, et al.
Pubblicazione: (2026) -
On the Limitations of Vision-Language Models in Understanding Image Transforms
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)