Universal Adversarial Attack on Aligned Multimodal LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Rahmatullaev, Temurbek, Druzhinina, Polina, Kurdiukov, Nikita, Mikhalchuk, Matvey, Kuznetsov, Andrey, Razzhigaev, Anton |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OmniFusion Technical Report
di: Goncharova, Elizaveta, et al.
Pubblicazione: (2024)
di: Goncharova, Elizaveta, et al.
Pubblicazione: (2024)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
di: Asanuma, Haruka, et al.
Pubblicazione: (2025)
di: Asanuma, Haruka, et al.
Pubblicazione: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
di: Skripkin, Matvey, et al.
Pubblicazione: (2025)
di: Skripkin, Matvey, et al.
Pubblicazione: (2025)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
di: Le, Van-Truong
Pubblicazione: (2026)
di: Le, Van-Truong
Pubblicazione: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
di: Simões, Lucca Emmanuel Pineli, et al.
Pubblicazione: (2024)
di: Simões, Lucca Emmanuel Pineli, et al.
Pubblicazione: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
di: Teo, Charlton
Pubblicazione: (2025)
di: Teo, Charlton
Pubblicazione: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
di: Bian, Zhipeng, et al.
Pubblicazione: (2025)
di: Bian, Zhipeng, et al.
Pubblicazione: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
di: Mulrooney, Alex, et al.
Pubblicazione: (2025)
di: Mulrooney, Alex, et al.
Pubblicazione: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
di: Ye, Hua, et al.
Pubblicazione: (2025)
di: Ye, Hua, et al.
Pubblicazione: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
di: Li, Danyang, et al.
Pubblicazione: (2025)
di: Li, Danyang, et al.
Pubblicazione: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
Defending against Backdoor Attacks via Module Switching
di: Li, Weijun, et al.
Pubblicazione: (2025)
di: Li, Weijun, et al.
Pubblicazione: (2025)
Leum-VL Technical Report
di: He, Yuxuan, et al.
Pubblicazione: (2026)
di: He, Yuxuan, et al.
Pubblicazione: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
Learning the meanings of function words from grounded language using a visual question answering model
di: Portelance, Eva, et al.
Pubblicazione: (2023)
di: Portelance, Eva, et al.
Pubblicazione: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Survey Transfer Learning: Recycling Data with Silicon Responses
di: Amini, Ali
Pubblicazione: (2025)
di: Amini, Ali
Pubblicazione: (2025)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
di: Marmoret, Axel, et al.
Pubblicazione: (2025)
di: Marmoret, Axel, et al.
Pubblicazione: (2025)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
di: Deiseroth, Björn, et al.
Pubblicazione: (2025)
di: Deiseroth, Björn, et al.
Pubblicazione: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2024)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2024)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
di: Sanders, Kate, et al.
Pubblicazione: (2024)
di: Sanders, Kate, et al.
Pubblicazione: (2024)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
di: Boumber, Dainis, et al.
Pubblicazione: (2024)
di: Boumber, Dainis, et al.
Pubblicazione: (2024)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
di: Agia, Christopher, et al.
Pubblicazione: (2024)
di: Agia, Christopher, et al.
Pubblicazione: (2024)
Documenti analoghi
-
OmniFusion Technical Report
di: Goncharova, Elizaveta, et al.
Pubblicazione: (2024) -
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
di: Asanuma, Haruka, et al.
Pubblicazione: (2025) -
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
di: Zhang, Sinin, et al.
Pubblicazione: (2026) -
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
di: Skripkin, Matvey, et al.
Pubblicazione: (2025) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
di: Le, Van-Truong
Pubblicazione: (2026)