Benchmarking Gaslighting Negation Attacks Against Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Bin, Yin, Hailong, Chen, Jingjing, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
von: Jiao, Pengkun, et al.
Veröffentlicht: (2025)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2025)
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
From Canteen Food to Daily Meals: Generalizing Food Recognition to More Practical Scenarios
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2024)
von: Li, Yian, et al.
Veröffentlicht: (2024)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
von: Zhao, Yunhan, et al.
Veröffentlicht: (2024)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2024)
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
von: Shah, Arya, et al.
Veröffentlicht: (2026)
von: Shah, Arya, et al.
Veröffentlicht: (2026)
White-box Multimodal Jailbreaks Against Large Vision-Language Models
von: Wang, Ruofan, et al.
Veröffentlicht: (2024)
von: Wang, Ruofan, et al.
Veröffentlicht: (2024)
Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
von: Deng, Linger, et al.
Veröffentlicht: (2024)
von: Deng, Linger, et al.
Veröffentlicht: (2024)
Model Inversion Attack Against Deep Hashing
von: Zhao, Dongdong, et al.
Veröffentlicht: (2025)
von: Zhao, Dongdong, et al.
Veröffentlicht: (2025)
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
von: Zhang, Ninghao, et al.
Veröffentlicht: (2026)
von: Zhang, Ninghao, et al.
Veröffentlicht: (2026)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
Negative Prototypes Guided Contrastive Learning for WSOD
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
Fake-HR1: Rethinking Reasoning of Vision Language Model for Synthetic Image Detection
von: Jiang, Changjiang, et al.
Veröffentlicht: (2026)
von: Jiang, Changjiang, et al.
Veröffentlicht: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
von: Han, Xianjing, et al.
Veröffentlicht: (2026)
von: Han, Xianjing, et al.
Veröffentlicht: (2026)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
Efficient Model-Based Purification Against Adversarial Attacks for LiDAR Segmentation
von: Gkillas, Alexandros, et al.
Veröffentlicht: (2025)
von: Gkillas, Alexandros, et al.
Veröffentlicht: (2025)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
von: Wang, Kangkang, et al.
Veröffentlicht: (2026)
von: Wang, Kangkang, et al.
Veröffentlicht: (2026)
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
Retrieval Augmented Recipe Generation
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
von: Du, Wanjun, et al.
Veröffentlicht: (2026)
von: Du, Wanjun, et al.
Veröffentlicht: (2026)
IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
von: Wang, Ruofan, et al.
Veröffentlicht: (2024)
von: Wang, Ruofan, et al.
Veröffentlicht: (2024)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
Tex3D: Objects as Attack Surfaces via Adversarial 3D Textures for Vision-Language-Action Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2026)
von: Chen, Jiawei, et al.
Veröffentlicht: (2026)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
von: Peng, Xingkai, et al.
Veröffentlicht: (2025)
von: Peng, Xingkai, et al.
Veröffentlicht: (2025)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
von: Hufe, Lorenz, et al.
Veröffentlicht: (2025)
von: Hufe, Lorenz, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
von: Jiao, Pengkun, et al.
Veröffentlicht: (2025) -
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026) -
From Canteen Food to Daily Meals: Generalizing Food Recognition to More Practical Scenarios
von: Liu, Guoshan, et al.
Veröffentlicht: (2024) -
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2024) -
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)