InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xunguang, Ji, Zhenlan, Ma, Pingchuan, Li, Zongjie, Wang, Shuai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
por: Yang, Shuai, et al.
Publicado: (2025)
por: Yang, Shuai, et al.
Publicado: (2025)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2025)
por: Wang, Xunguang, et al.
Publicado: (2025)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
por: Gan, Yulu, et al.
Publicado: (2023)
por: Gan, Yulu, et al.
Publicado: (2023)
Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following
por: Kang, Myeongkyun, et al.
Publicado: (2026)
por: Kang, Myeongkyun, et al.
Publicado: (2026)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
por: Zhang, Jinrui, et al.
Publicado: (2024)
por: Zhang, Jinrui, et al.
Publicado: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
por: Cao, Meng, et al.
Publicado: (2024)
por: Cao, Meng, et al.
Publicado: (2024)
MTAttack: Multi-Target Backdoor Attacks against Large Vision-Language Models
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
por: Ma, Ji, et al.
Publicado: (2024)
por: Ma, Ji, et al.
Publicado: (2024)
Learning to Instruct for Visual Instruction Tuning
por: Zhou, Zhihan, et al.
Publicado: (2025)
por: Zhou, Zhihan, et al.
Publicado: (2025)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2026)
por: Wang, Xunguang, et al.
Publicado: (2026)
Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
por: Cao, Yiming, et al.
Publicado: (2025)
por: Cao, Yiming, et al.
Publicado: (2025)
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
por: Zhang, Wenqi, et al.
Publicado: (2024)
por: Zhang, Wenqi, et al.
Publicado: (2024)
Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
por: Hu, Rui, et al.
Publicado: (2024)
por: Hu, Rui, et al.
Publicado: (2024)
SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models
por: Wang, Haobo, et al.
Publicado: (2026)
por: Wang, Haobo, et al.
Publicado: (2026)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
por: Park, Dongmin, et al.
Publicado: (2024)
por: Park, Dongmin, et al.
Publicado: (2024)
Multimodal Instruction Tuning with Hybrid State Space Models
por: Zhou, Jianing, et al.
Publicado: (2024)
por: Zhou, Jianing, et al.
Publicado: (2024)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
por: Wei, Cong, et al.
Publicado: (2024)
por: Wei, Cong, et al.
Publicado: (2024)
DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
por: Chen, Yangyi, et al.
Publicado: (2023)
por: Chen, Yangyi, et al.
Publicado: (2023)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2025)
por: Wang, Xunguang, et al.
Publicado: (2025)
Improving Transferable Targeted Attacks with Feature Tuning Mixup
por: Liang, Kaisheng, et al.
Publicado: (2024)
por: Liang, Kaisheng, et al.
Publicado: (2024)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
por: Li, Juncheng, et al.
Publicado: (2025)
por: Li, Juncheng, et al.
Publicado: (2025)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
por: Mei, Hefei, et al.
Publicado: (2025)
por: Mei, Hefei, et al.
Publicado: (2025)
Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
por: Ma, Ji, et al.
Publicado: (2025)
por: Ma, Ji, et al.
Publicado: (2025)
Visual Instruction Tuning with Chain of Region-of-Interest
por: Chen, Yixin, et al.
Publicado: (2025)
por: Chen, Yixin, et al.
Publicado: (2025)
SEAGULL: No-reference Image Quality Assessment for Regions of Interest via Vision-Language Instruction Tuning
por: Chen, Zewen, et al.
Publicado: (2024)
por: Chen, Zewen, et al.
Publicado: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
por: Wang, Zeyu, et al.
Publicado: (2025)
por: Wang, Zeyu, et al.
Publicado: (2025)
InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
por: Zhao, Ruoyu, et al.
Publicado: (2024)
por: Zhao, Ruoyu, et al.
Publicado: (2024)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
por: Liao, Ning, et al.
Publicado: (2023)
por: Liao, Ning, et al.
Publicado: (2023)
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
por: Zhan, Yang, et al.
Publicado: (2024)
por: Zhan, Yang, et al.
Publicado: (2024)
Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
por: Wang, Ruofan, et al.
Publicado: (2025)
por: Wang, Ruofan, et al.
Publicado: (2025)
InstructSAM: Segment Any Instance with Any Instructions
por: Yuan, Yuqian, et al.
Publicado: (2026)
por: Yuan, Yuqian, et al.
Publicado: (2026)
InstructRestore: Region-Customized Image Restoration with Human Instructions
por: Liu, Shuaizheng, et al.
Publicado: (2025)
por: Liu, Shuaizheng, et al.
Publicado: (2025)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
por: Jiang, Lei, et al.
Publicado: (2025)
por: Jiang, Lei, et al.
Publicado: (2025)
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
por: Wang, Wenqing, et al.
Publicado: (2026)
por: Wang, Wenqing, et al.
Publicado: (2026)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
por: Liu, Jihao, et al.
Publicado: (2024)
por: Liu, Jihao, et al.
Publicado: (2024)
InstructEngine: Instruction-driven Text-to-Image Alignment
por: Lu, Xingyu, et al.
Publicado: (2025)
por: Lu, Xingyu, et al.
Publicado: (2025)
Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
por: Peng, Jihua, et al.
Publicado: (2025)
por: Peng, Jihua, et al.
Publicado: (2025)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
por: Yan, Qiao, et al.
Publicado: (2025)
por: Yan, Qiao, et al.
Publicado: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
por: Li, Bin, et al.
Publicado: (2025)
por: Li, Bin, et al.
Publicado: (2025)
InstructVEdit: A Holistic Approach for Instructional Video Editing
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
Ejemplares similares
-
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
por: Yang, Shuai, et al.
Publicado: (2025) -
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2025) -
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
por: Gan, Yulu, et al.
Publicado: (2023) -
Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following
por: Kang, Myeongkyun, et al.
Publicado: (2026) -
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
por: Zhang, Jinrui, et al.
Publicado: (2024)