VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Bowen, Xiang, Yongli, Hong, Ziming, Lin, Zerong, Yu, Chaojian, Liu, Tongliang, You, Xinge |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking the Non-Transferable Barrier via Test-Time Data Disguising
by: Xiang, Yongli, et al.
Published: (2025)
by: Xiang, Yongli, et al.
Published: (2025)
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval
by: Xu, Yizhuo, et al.
Published: (2026)
by: Xu, Yizhuo, et al.
Published: (2026)
Toward Robust Non-Transferable Learning: A Survey and Benchmark
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance
by: Xiang, Yongli, et al.
Published: (2026)
by: Xiang, Yongli, et al.
Published: (2026)
Prototype-Guided Curriculum Learning for Zero-Shot Learning
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning
by: Hou, Wenjin, et al.
Published: (2024)
by: Hou, Wenjin, et al.
Published: (2024)
RunawayEvil: Jailbreaking the Image-to-Video Generative Models
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Detail Reinforcement Diffusion Model: Augmentation Fine-Grained Visual Categorization in Few-Shot Conditions
by: Wu, Tianxu, et al.
Published: (2023)
by: Wu, Tianxu, et al.
Published: (2023)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
by: Pu, Yujiang, et al.
Published: (2025)
by: Pu, Yujiang, et al.
Published: (2025)
Discriminative Image Generation with Diffusion Models for Zero-Shot Learning
by: Fu, Dingjie, et al.
Published: (2024)
by: Fu, Dingjie, et al.
Published: (2024)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps
by: Zhang, Jianwei, et al.
Published: (2026)
by: Zhang, Jianwei, et al.
Published: (2026)
Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching
by: Lin, Yexiong, et al.
Published: (2025)
by: Lin, Yexiong, et al.
Published: (2025)
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025)
by: Lin, Gaojie, et al.
Published: (2025)
Early-Bird Diffusion: Investigating and Leveraging Timestep-Aware Early-Bird Tickets in Diffusion Models for Efficient Training
by: Whalen, Lexington, et al.
Published: (2025)
by: Whalen, Lexington, et al.
Published: (2025)
Encore: Conditioning Trajectory Forecasting via Biased Ego Rehearsals
by: Wong, Conghao, et al.
Published: (2026)
by: Wong, Conghao, et al.
Published: (2026)
Omni-Recon: Harnessing Image-based Rendering for General-Purpose Neural Radiance Fields
by: Fu, Yonggan, et al.
Published: (2024)
by: Fu, Yonggan, et al.
Published: (2024)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
by: Qu, Qiang, et al.
Published: (2025)
by: Qu, Qiang, et al.
Published: (2025)
AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos
by: Chen, Yushuo, et al.
Published: (2024)
by: Chen, Yushuo, et al.
Published: (2024)
In-Video Instructions: Visual Signals as Generative Control
by: Fang, Gongfan, et al.
Published: (2025)
by: Fang, Gongfan, et al.
Published: (2025)
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
by: Yu, Mingyu, et al.
Published: (2026)
by: Yu, Mingyu, et al.
Published: (2026)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026)
by: Zhang, Yabo, et al.
Published: (2026)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
by: Jiang, Huajie, et al.
Published: (2025)
by: Jiang, Huajie, et al.
Published: (2025)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
by: Zhang, Yabo, et al.
Published: (2024)
by: Zhang, Yabo, et al.
Published: (2024)
RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing
by: Zhao, Longjie, et al.
Published: (2025)
by: Zhao, Longjie, et al.
Published: (2025)
Intellectual Property Protection for 3D Gaussian Splatting Assets: A Survey
by: Zhao, Longjie, et al.
Published: (2026)
by: Zhao, Longjie, et al.
Published: (2026)
SocialCircle+: Learning the Angle-based Conditioned Interaction Representation for Pedestrian Trajectory Prediction
by: Wong, Conghao, et al.
Published: (2024)
by: Wong, Conghao, et al.
Published: (2024)
Reverberation: Learning the Latencies Before Forecasting Trajectories
by: Wong, Conghao, et al.
Published: (2025)
by: Wong, Conghao, et al.
Published: (2025)
Resonance: Learning to Predict Social-Aware Pedestrian Trajectories as Co-Vibrations
by: Wong, Conghao, et al.
Published: (2024)
by: Wong, Conghao, et al.
Published: (2024)
What Happens Without Background? Constructing Foreground-Only Data for Fine-Grained Tasks
by: Wang, Yuetian, et al.
Published: (2024)
by: Wang, Yuetian, et al.
Published: (2024)
Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
by: Huang, Zhuo, et al.
Published: (2023)
by: Huang, Zhuo, et al.
Published: (2023)
FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing
by: Li, Maomao, et al.
Published: (2026)
by: Li, Maomao, et al.
Published: (2026)
Diffusion Models Are Innate One-Step Generators
by: Zheng, Bowen, et al.
Published: (2024)
by: Zheng, Bowen, et al.
Published: (2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
VideoNeuMat: Neural Material Extraction from Generative Video Models
by: Xue, Bowen, et al.
Published: (2026)
by: Xue, Bowen, et al.
Published: (2026)
Similar Items
-
Jailbreaking the Non-Transferable Barrier via Test-Time Data Disguising
by: Xiang, Yongli, et al.
Published: (2025) -
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval
by: Xu, Yizhuo, et al.
Published: (2026) -
Toward Robust Non-Transferable Learning: A Survey and Benchmark
by: Hong, Ziming, et al.
Published: (2025) -
When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance
by: Xiang, Yongli, et al.
Published: (2026) -
Prototype-Guided Curriculum Learning for Zero-Shot Learning
by: Wang, Lei, et al.
Published: (2025)