Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xiaohan, Li, Zhaoyi, Luo, Yaxin, Cui, Jiacheng, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
by: Luo, Yaxin, et al.
Published: (2025)
by: Luo, Yaxin, et al.
Published: (2025)
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
by: Li, Zhaoyi, et al.
Published: (2025)
by: Li, Zhaoyi, et al.
Published: (2025)
Dataset Distillation via Committee Voting
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
by: Luo, Yaxin, et al.
Published: (2026)
by: Luo, Yaxin, et al.
Published: (2026)
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
by: Gordon, Brian, et al.
Published: (2025)
by: Gordon, Brian, et al.
Published: (2025)
Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention
by: Mei, Hefei, et al.
Published: (2026)
by: Mei, Hefei, et al.
Published: (2026)
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
by: Gong, Xuan, et al.
Published: (2024)
by: Gong, Xuan, et al.
Published: (2024)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
by: Liang, Zihan, et al.
Published: (2025)
by: Liang, Zihan, et al.
Published: (2025)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Generating Fine Details of Entity Interactions
by: Gu, Xinyi, et al.
Published: (2025)
by: Gu, Xinyi, et al.
Published: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
by: Xia, Runze, et al.
Published: (2025)
by: Xia, Runze, et al.
Published: (2025)
Exploring 3D Dataset Pruning
by: Zhao, Xiaohan, et al.
Published: (2026)
by: Zhao, Xiaohan, et al.
Published: (2026)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
by: Cui, Chenhang, et al.
Published: (2024)
by: Cui, Chenhang, et al.
Published: (2024)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
by: Wei, Canshi
Published: (2024)
by: Wei, Canshi
Published: (2024)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025)
by: An, Zhaoyi, et al.
Published: (2025)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding
by: Guo, Leilei, et al.
Published: (2025)
by: Guo, Leilei, et al.
Published: (2025)
Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
by: Nguyen, Tuan Dung, et al.
Published: (2026)
by: Nguyen, Tuan Dung, et al.
Published: (2026)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
by: Chen, Wenting, et al.
Published: (2023)
by: Chen, Wenting, et al.
Published: (2023)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
by: Park, Seongheon, et al.
Published: (2026)
by: Park, Seongheon, et al.
Published: (2026)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
by: Liu, Junzhuo, et al.
Published: (2024)
by: Liu, Junzhuo, et al.
Published: (2024)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
by: Deng, Haolin, et al.
Published: (2026)
by: Deng, Haolin, et al.
Published: (2026)
GreedyPixel: Fine-Grained Black-Box Adversarial Attack Via Greedy Algorithm
by: Wang, Hanrui, et al.
Published: (2025)
by: Wang, Hanrui, et al.
Published: (2025)
STEMTOX: From Social Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning
by: Swain, Subhankar, et al.
Published: (2025)
by: Swain, Subhankar, et al.
Published: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
by: Asokan, Mothilal, et al.
Published: (2025)
by: Asokan, Mothilal, et al.
Published: (2025)
SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
by: Zinkovich, Viktoriia, et al.
Published: (2025)
by: Zinkovich, Viktoriia, et al.
Published: (2025)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
by: Wang, Sudong, et al.
Published: (2026)
by: Wang, Sudong, et al.
Published: (2026)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
by: Qu, Renyi, et al.
Published: (2024)
by: Qu, Renyi, et al.
Published: (2024)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
by: Wang, Yeyuan, et al.
Published: (2025)
by: Wang, Yeyuan, et al.
Published: (2025)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
by: Luo, Gen, et al.
Published: (2024)
by: Luo, Gen, et al.
Published: (2024)
Low-Frequency Black-Box Backdoor Attack via Evolutionary Algorithm
by: Qiao, Yanqi, et al.
Published: (2024)
by: Qiao, Yanqi, et al.
Published: (2024)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
by: Liu, Runzhou, et al.
Published: (2026)
by: Liu, Runzhou, et al.
Published: (2026)
Similar Items
-
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
by: Luo, Yaxin, et al.
Published: (2025) -
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
by: Li, Zhaoyi, et al.
Published: (2025) -
Dataset Distillation via Committee Voting
by: Cui, Jiacheng, et al.
Published: (2025) -
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
by: Luo, Yaxin, et al.
Published: (2026) -
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
by: Cui, Jiacheng, et al.
Published: (2025)