Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Schaeffer, Rylan, Valentine, Dan, Bailey, Luke, Chua, James, Eyzaguirre, Cristóbal, Durante, Zane, Benton, Joe, Miranda, Brando, Sleight, Henry, Hughes, John, Agrawal, Rajashree, Sharma, Mrinank, Emmons, Scott, Koyejo, Sanmi, Perez, Ethan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024)
by: Wang, Tony T., et al.
Published: (2024)
Pretraining Scaling Laws for Generative Evaluations of Language Models
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
In-Context Learning of Energy Functions
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)
by: Obbad, Elyas, et al.
Published: (2024)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
by: Peng, Alwin, et al.
Published: (2024)
by: Peng, Alwin, et al.
Published: (2024)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
by: Gerstgrasser, Matthias, et al.
Published: (2024)
by: Gerstgrasser, Matthias, et al.
Published: (2024)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
by: Gupta, Isha, et al.
Published: (2025)
by: Gupta, Isha, et al.
Published: (2025)
What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes
by: Lecomte, Victor, et al.
Published: (2023)
by: Lecomte, Victor, et al.
Published: (2023)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Chain-of-Thought Hijacking
by: Zhao, Jianli, et al.
Published: (2025)
by: Zhao, Jianli, et al.
Published: (2025)
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Investigating Data Contamination for Pre-training Language Models
by: Jiang, Minhao, et al.
Published: (2024)
by: Jiang, Minhao, et al.
Published: (2024)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
by: Kazdan, Joshua, et al.
Published: (2024)
by: Kazdan, Joshua, et al.
Published: (2024)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
by: Denisov-Blanch, Yegor, et al.
Published: (2026)
by: Denisov-Blanch, Yegor, et al.
Published: (2026)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
How Do Large Language Monkeys Get Their Power (Laws)?
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
by: Schaeffer, Rylan, et al.
Published: (2026)
by: Schaeffer, Rylan, et al.
Published: (2026)
Pantograph: A Machine-to-Machine Interaction Interface for Advanced Theorem Proving, High Level Reasoning, and Data Extraction in Lean 4
by: Aniva, Leni, et al.
Published: (2024)
by: Aniva, Leni, et al.
Published: (2024)
Quantifying Variance in Evaluation Benchmarks
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
by: Youstra, Jack, et al.
Published: (2025)
by: Youstra, Jack, et al.
Published: (2025)
Scale Dependent Data Duplication
by: Kazdan, Joshua, et al.
Published: (2026)
by: Kazdan, Joshua, et al.
Published: (2026)
Is Pre-training Truly Better Than Meta-Learning?
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
by: Robertson, Zachary, et al.
Published: (2025)
by: Robertson, Zachary, et al.
Published: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
EXPANSIÓN Y LÍMITES DE LA BUENA FE OBJETIVA – A PROPÓSITO DEL “PROYECTO DE PRINCIPIOS LATINOAMERICANOS DE DERECHO DE LOS CONTRATOS”
by: Cristóbal Eyzaguirre Baeza
Published: (2013)
by: Cristóbal Eyzaguirre Baeza
Published: (2013)
A Framework for Objective-Driven Dynamical Stochastic Fields
by: Zhang, Yibo Jacky, et al.
Published: (2025)
by: Zhang, Yibo Jacky, et al.
Published: (2025)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023)
by: Bailey, Luke, et al.
Published: (2023)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025)
by: Chan, Willy, et al.
Published: (2025)
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
by: McGuinness, Max, et al.
Published: (2025)
by: McGuinness, Max, et al.
Published: (2025)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
High-Dimensional Markov-switching Ordinary Differential Processes
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
Distributional Machine Unlearning via Selective Data Removal
by: Allouah, Youssef, et al.
Published: (2025)
by: Allouah, Youssef, et al.
Published: (2025)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
by: Zhu, Junzhe, et al.
Published: (2023)
by: Zhu, Junzhe, et al.
Published: (2023)
Similar Items
-
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024) -
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024) -
Pretraining Scaling Laws for Generative Evaluations of Language Models
by: Schaeffer, Rylan, et al.
Published: (2025) -
In-Context Learning of Energy Functions
by: Schaeffer, Rylan, et al.
Published: (2024) -
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)