AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yuqi, Miao, Yuchun, Li, Zuchao, Ding, Liang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Intention Analysis Makes LLMs A Good Jailbreak Defender
di: Zhang, Yuqi, et al.
Pubblicazione: (2024)
di: Zhang, Yuqi, et al.
Pubblicazione: (2024)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
di: Zhou, Andy, et al.
Pubblicazione: (2024)
di: Zhou, Andy, et al.
Pubblicazione: (2024)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
di: Li, Yanshu, et al.
Pubblicazione: (2025)
di: Li, Yanshu, et al.
Pubblicazione: (2025)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
di: Miao, Ziqi, et al.
Pubblicazione: (2025)
di: Miao, Ziqi, et al.
Pubblicazione: (2025)
Improving Alignment in LVLMs with Debiased Self-Judgment
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
di: Zheng, Haohan, et al.
Pubblicazione: (2025)
di: Zheng, Haohan, et al.
Pubblicazione: (2025)
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
di: He, Jinghan, et al.
Pubblicazione: (2024)
di: He, Jinghan, et al.
Pubblicazione: (2024)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
di: Yuan, Fan, et al.
Pubblicazione: (2024)
di: Yuan, Fan, et al.
Pubblicazione: (2024)
VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing
di: Huang, Yanbin, et al.
Pubblicazione: (2026)
di: Huang, Yanbin, et al.
Pubblicazione: (2026)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
di: Zhang, Xiaofeng, et al.
Pubblicazione: (2024)
di: Zhang, Xiaofeng, et al.
Pubblicazione: (2024)
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2023)
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2023)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
di: Lu, Liming, et al.
Pubblicazione: (2025)
di: Lu, Liming, et al.
Pubblicazione: (2025)
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
di: Liu, Xunzhuo, et al.
Pubblicazione: (2026)
di: Liu, Xunzhuo, et al.
Pubblicazione: (2026)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
di: Fang, Hao, et al.
Pubblicazione: (2025)
di: Fang, Hao, et al.
Pubblicazione: (2025)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
di: Jiang, Lei, et al.
Pubblicazione: (2025)
di: Jiang, Lei, et al.
Pubblicazione: (2025)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
di: Wei, Zhihua, et al.
Pubblicazione: (2026)
di: Wei, Zhihua, et al.
Pubblicazione: (2026)
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
di: Zheng, Ge, et al.
Pubblicazione: (2025)
di: Zheng, Ge, et al.
Pubblicazione: (2025)
Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
di: Guo, Xin, et al.
Pubblicazione: (2025)
di: Guo, Xin, et al.
Pubblicazione: (2025)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
di: Li, Yifan, et al.
Pubblicazione: (2024)
di: Li, Yifan, et al.
Pubblicazione: (2024)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
di: Zheng, Huan, et al.
Pubblicazione: (2025)
di: Zheng, Huan, et al.
Pubblicazione: (2025)
TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks
di: Zou, Quanchen, et al.
Pubblicazione: (2026)
di: Zou, Quanchen, et al.
Pubblicazione: (2026)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
di: Leng, Jixuan, et al.
Pubblicazione: (2025)
di: Leng, Jixuan, et al.
Pubblicazione: (2025)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
di: Salamatian, Ali, et al.
Pubblicazione: (2025)
di: Salamatian, Ali, et al.
Pubblicazione: (2025)
A Multi-view Mask Contrastive Learning Graph Convolutional Neural Network for Age Estimation
di: Zhang, Yiping, et al.
Pubblicazione: (2024)
di: Zhang, Yiping, et al.
Pubblicazione: (2024)
Centered Masking for Language-Image Pre-Training
di: Liang, Mingliang, et al.
Pubblicazione: (2024)
di: Liang, Mingliang, et al.
Pubblicazione: (2024)
Robust Anti-Backdoor Instruction Tuning in LVLMs
di: Xun, Yuan, et al.
Pubblicazione: (2025)
di: Xun, Yuan, et al.
Pubblicazione: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
di: Tang, Zicong, et al.
Pubblicazione: (2025)
di: Tang, Zicong, et al.
Pubblicazione: (2025)
Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
di: Sui, Xiangjie, et al.
Pubblicazione: (2025)
di: Sui, Xiangjie, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
di: Zhou, Qi, et al.
Pubblicazione: (2024)
di: Zhou, Qi, et al.
Pubblicazione: (2024)
Breaking through Deterministic Barriers: Randomized Pruning Mask Generation and Selection
di: Li, Jianwei, et al.
Pubblicazione: (2023)
di: Li, Jianwei, et al.
Pubblicazione: (2023)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
di: Pi, Renjie, et al.
Pubblicazione: (2024)
di: Pi, Renjie, et al.
Pubblicazione: (2024)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Intention Analysis Makes LLMs A Good Jailbreak Defender
di: Zhang, Yuqi, et al.
Pubblicazione: (2024) -
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
di: Zhou, Andy, et al.
Pubblicazione: (2024) -
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
di: Li, Yanshu, et al.
Pubblicazione: (2025) -
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
di: Miao, Ziqi, et al.
Pubblicazione: (2025) -
Improving Alignment in LVLMs with Debiased Self-Judgment
di: Yang, Sihan, et al.
Pubblicazione: (2025)