On Optimizing Multimodal Jailbreaks for Spoken Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Krishnan, Aravind, Stańczak, Karolina, Klakow, Dietrich |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Block-Operations: Using Modular Routing to Improve Compositional Generalization
por: Dietz, Florian, et al.
Publicado: (2024)
por: Dietz, Florian, et al.
Publicado: (2024)
IGC: Integrating a Gated Calculator into an LLM to Solve Arithmetic Tasks Reliably and Efficiently
por: Dietz, Florian, et al.
Publicado: (2025)
por: Dietz, Florian, et al.
Publicado: (2025)
Chemical Language Models for Natural Products: A State-Space Model Approach
por: Wang, Ho-Hsuan, et al.
Publicado: (2026)
por: Wang, Ho-Hsuan, et al.
Publicado: (2026)
Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
por: Blatt, Alexander, et al.
Publicado: (2024)
por: Blatt, Alexander, et al.
Publicado: (2024)
On the Encoding of Gender in Transformer-based ASR Representations
por: Krishnan, Aravind, et al.
Publicado: (2024)
por: Krishnan, Aravind, et al.
Publicado: (2024)
Comgra: A Tool for Analyzing and Debugging Neural Networks
por: Dietz, Florian, et al.
Publicado: (2024)
por: Dietz, Florian, et al.
Publicado: (2024)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
por: García-de-Herreros, Paloma, et al.
Publicado: (2025)
por: García-de-Herreros, Paloma, et al.
Publicado: (2025)
Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters
por: Chen, Zixuan, et al.
Publicado: (2025)
por: Chen, Zixuan, et al.
Publicado: (2025)
Jailbreaking Attack against Multimodal Large Language Model
por: Niu, Zhenxing, et al.
Publicado: (2024)
por: Niu, Zhenxing, et al.
Publicado: (2024)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
por: Jia, Xiaojun, et al.
Publicado: (2024)
por: Jia, Xiaojun, et al.
Publicado: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024)
por: Chao, Patrick, et al.
Publicado: (2024)
TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
por: Xiong, Xiqiao, et al.
Publicado: (2025)
por: Xiong, Xiqiao, et al.
Publicado: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
por: Ramapuram, Jason, et al.
Publicado: (2026)
por: Ramapuram, Jason, et al.
Publicado: (2026)
Aligning Pre-trained Models for Spoken Language Translation
por: Sedláček, Šimon, et al.
Publicado: (2024)
por: Sedláček, Šimon, et al.
Publicado: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
por: Oh, Sejoon, et al.
Publicado: (2024)
por: Oh, Sejoon, et al.
Publicado: (2024)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
por: Geng, Jiahui, et al.
Publicado: (2025)
por: Geng, Jiahui, et al.
Publicado: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
por: Lee, Isack, et al.
Publicado: (2024)
por: Lee, Isack, et al.
Publicado: (2024)
Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting
por: Gundluru, Ramesh, et al.
Publicado: (2025)
por: Gundluru, Ramesh, et al.
Publicado: (2025)
Jailbreaking Large Language Models in Infinitely Many Ways
por: Goldstein, Oliver, et al.
Publicado: (2025)
por: Goldstein, Oliver, et al.
Publicado: (2025)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
por: Li, Xuan, et al.
Publicado: (2023)
por: Li, Xuan, et al.
Publicado: (2023)
JULI: Jailbreak Large Language Models by Self-Introspection
por: Wang, Jesson, et al.
Publicado: (2025)
por: Wang, Jesson, et al.
Publicado: (2025)
Value Drifts: Tracing Value Alignment During LLM Post-Training
por: Bhatia, Mehar, et al.
Publicado: (2025)
por: Bhatia, Mehar, et al.
Publicado: (2025)
Transformers for molecular property prediction: Domain adaptation efficiently improves performance
por: Sultan, Afnan, et al.
Publicado: (2025)
por: Sultan, Afnan, et al.
Publicado: (2025)
The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?
por: Kao, Ching-Chia, et al.
Publicado: (2024)
por: Kao, Ching-Chia, et al.
Publicado: (2024)
GuardNet: Graph-Attention Filtering for Jailbreak Defense in Large Language Models
por: Forough, Javad, et al.
Publicado: (2025)
por: Forough, Javad, et al.
Publicado: (2025)
Multimodal Pragmatic Jailbreak on Text-to-image Models
por: Liu, Tong, et al.
Publicado: (2024)
por: Liu, Tong, et al.
Publicado: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024)
por: Peng, Benji, et al.
Publicado: (2024)
A Multilingual Perspective on Probing Gender Bias
por: Stańczak, Karolina
Publicado: (2024)
por: Stańczak, Karolina
Publicado: (2024)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
por: Wang, Xiangwen, et al.
Publicado: (2026)
por: Wang, Xiangwen, et al.
Publicado: (2026)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
por: Li, Songze, et al.
Publicado: (2026)
por: Li, Songze, et al.
Publicado: (2026)
Jailbreaking Black Box Large Language Models in Twenty Queries
por: Chao, Patrick, et al.
Publicado: (2023)
por: Chao, Patrick, et al.
Publicado: (2023)
Spoken Language Intelligence of Large Language Models for Language Learning
por: Peng, Linkai, et al.
Publicado: (2023)
por: Peng, Linkai, et al.
Publicado: (2023)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
por: Zhou, Andy, et al.
Publicado: (2024)
por: Zhou, Andy, et al.
Publicado: (2024)
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
por: Yang, Xikang, et al.
Publicado: (2024)
por: Yang, Xikang, et al.
Publicado: (2024)
Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
por: Wang, Zhaoxin, et al.
Publicado: (2025)
por: Wang, Zhaoxin, et al.
Publicado: (2025)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
por: Zhang, Ziyi, et al.
Publicado: (2025)
por: Zhang, Ziyi, et al.
Publicado: (2025)
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
por: Li, Yuxi, et al.
Publicado: (2024)
por: Li, Yuxi, et al.
Publicado: (2024)
Learning Program Behavioral Models from Synthesized Input-Output Pairs
por: Mammadov, Tural, et al.
Publicado: (2024)
por: Mammadov, Tural, et al.
Publicado: (2024)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
por: Jin, Haibo, et al.
Publicado: (2024)
por: Jin, Haibo, et al.
Publicado: (2024)
Classifier Language Models: Unifying Sparse Finetuning and Adaptive Tokenization for Specialized Classification Tasks
por: Krishnan, Adit, et al.
Publicado: (2025)
por: Krishnan, Adit, et al.
Publicado: (2025)
Ejemplares similares
-
Block-Operations: Using Modular Routing to Improve Compositional Generalization
por: Dietz, Florian, et al.
Publicado: (2024) -
IGC: Integrating a Gated Calculator into an LLM to Solve Arithmetic Tasks Reliably and Efficiently
por: Dietz, Florian, et al.
Publicado: (2025) -
Chemical Language Models for Natural Products: A State-Space Model Approach
por: Wang, Ho-Hsuan, et al.
Publicado: (2026) -
Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
por: Blatt, Alexander, et al.
Publicado: (2024) -
On the Encoding of Gender in Transformer-based ASR Representations
por: Krishnan, Aravind, et al.
Publicado: (2024)