When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
Fuente:
arXiv
Salvato in:
| Autori principali: | Dingeto, Hiskias, Kwon, Taeyoun, Choi, Dasol, Kim, Bodam, Lee, DongGeon, Park, Haon, Lee, JaeHoon, Shin, Jongho |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025)
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025)
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
di: Dingeto, Hiskias, et al.
Pubblicazione: (2026)
di: Dingeto, Hiskias, et al.
Pubblicazione: (2026)
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
di: Choi, Dasol, et al.
Pubblicazione: (2026)
di: Choi, Dasol, et al.
Pubblicazione: (2026)
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
di: Kim, Kihyun, et al.
Pubblicazione: (2026)
di: Kim, Kihyun, et al.
Pubblicazione: (2026)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models
di: Choi, Dasol, et al.
Pubblicazione: (2026)
di: Choi, Dasol, et al.
Pubblicazione: (2026)
Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark
di: Jeong, Jihae, et al.
Pubblicazione: (2025)
di: Jeong, Jihae, et al.
Pubblicazione: (2025)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
ToDi: Token-wise Distillation via Fine-Grained Divergence Control
di: Jung, Seongryong, et al.
Pubblicazione: (2025)
di: Jung, Seongryong, et al.
Pubblicazione: (2025)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
Typed-RAG: Type-Aware Decomposition of Non-Factoid Questions for Retrieval-Augmented Generation
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
Infinite-level Fock spaces, crystal bases, and tensor product of extremal weight modules of type $A_{+\infty}$
di: Kwon, Jae-Hoon, et al.
Pubblicazione: (2025)
di: Kwon, Jae-Hoon, et al.
Pubblicazione: (2025)
Theme-Explanation Structure for Table Summarization using Large Language Models: A Case Study on Korean Tabular Data
di: Kwack, TaeYoon, et al.
Pubblicazione: (2025)
di: Kwack, TaeYoon, et al.
Pubblicazione: (2025)
Learning While Transmitting: Pilotless Polar Coded Modulation for Short Packet Transmission
di: Choi, Geon, et al.
Pubblicazione: (2026)
di: Choi, Geon, et al.
Pubblicazione: (2026)
Rate-Matching Deep Polar Codes via Polar Coded Extension
di: Choi, Geon, et al.
Pubblicazione: (2025)
di: Choi, Geon, et al.
Pubblicazione: (2025)
Sparsely Pre-transformed Polar Codes for Low-Latency SCL Decoding
di: Choi, Geon, et al.
Pubblicazione: (2024)
di: Choi, Geon, et al.
Pubblicazione: (2024)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
MARIOH: Multiplicity-Aware Hypergraph Reconstruction
di: Lee, Kyuhan, et al.
Pubblicazione: (2025)
di: Lee, Kyuhan, et al.
Pubblicazione: (2025)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
di: Yan, Yu, et al.
Pubblicazione: (2025)
di: Yan, Yu, et al.
Pubblicazione: (2025)
Discrimination between ventricular tachycardia and wide‐QRS preexcited tachycardia
di: Jae Hoon Lee
Pubblicazione: (2024)
di: Jae Hoon Lee
Pubblicazione: (2024)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
di: Kim, Hee-Seon, et al.
Pubblicazione: (2025)
di: Kim, Hee-Seon, et al.
Pubblicazione: (2025)
ObjexMT: Objective Extraction and Metacognitive Calibration for LLM-as-a-Judge under Multi-Turn Jailbreaks
di: Kim, Hyunjun, et al.
Pubblicazione: (2025)
di: Kim, Hyunjun, et al.
Pubblicazione: (2025)
X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
di: Kim, Hyunjun, et al.
Pubblicazione: (2025)
di: Kim, Hyunjun, et al.
Pubblicazione: (2025)
Improvement of Glucose Metabolism by Pennogenin 3‐O‐β‐Chacotrioside via Activation of IRS/PI3K/Akt Signaling and Mitochondrial Respiration in Insulin‐Resistant Hepatocytes
di: Jae‐In Lee, et al.
Pubblicazione: (2025)
di: Jae‐In Lee, et al.
Pubblicazione: (2025)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
di: Lee, Dongheon, et al.
Pubblicazione: (2024)
di: Lee, Dongheon, et al.
Pubblicazione: (2024)
Oscillator representations of quantum affine orthosymplectic superalgebras
di: Kwon, Jae-Hoon, et al.
Pubblicazione: (2023)
di: Kwon, Jae-Hoon, et al.
Pubblicazione: (2023)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
di: Choi, Hahyeon, et al.
Pubblicazione: (2025)
di: Choi, Hahyeon, et al.
Pubblicazione: (2025)
Mine-JEPA: In-Domain Self-Supervised Learning for Mine-Like Object Classification in Side-Scan Sonar
di: Kwon, Taeyoun, et al.
Pubblicazione: (2026)
di: Kwon, Taeyoun, et al.
Pubblicazione: (2026)
Heuristic Algorithm-based Action Masking Reinforcement Learning (HAAM-RL) with Ensemble Inference Method
di: Choi, Kyuwon, et al.
Pubblicazione: (2024)
di: Choi, Kyuwon, et al.
Pubblicazione: (2024)
When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
di: Kim, Hanna, et al.
Pubblicazione: (2024)
di: Kim, Hanna, et al.
Pubblicazione: (2024)
A New Paradigm Integrating the Concepts of Particle Abrasion and Breakage
di: Tripathi, Priya, et al.
Pubblicazione: (2023)
di: Tripathi, Priya, et al.
Pubblicazione: (2023)
Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy
di: Kim, Juyeon, et al.
Pubblicazione: (2025)
di: Kim, Juyeon, et al.
Pubblicazione: (2025)
DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval
di: Park, Geon, et al.
Pubblicazione: (2026)
di: Park, Geon, et al.
Pubblicazione: (2026)
Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues
di: Kwon, Younghoo, et al.
Pubblicazione: (2025)
di: Kwon, Younghoo, et al.
Pubblicazione: (2025)
Double-pass rotating z-cut quartz plate as a rapidly variable waveplate
di: Lee, Byungjin, et al.
Pubblicazione: (2025)
di: Lee, Byungjin, et al.
Pubblicazione: (2025)
El impacto de la variable de género en la migración Honduras-México: el caso de las Hondureñas en Frontera Comalapa
di: Nicanor Madueño Haon
Pubblicazione: (2010)
di: Nicanor Madueño Haon
Pubblicazione: (2010)
Distribution-Level Feature Distancing for Machine Unlearning: Towards a Better Trade-off Between Model Utility and Forgetting
di: Choi, Dasol, et al.
Pubblicazione: (2024)
di: Choi, Dasol, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025) -
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
di: Dingeto, Hiskias, et al.
Pubblicazione: (2026) -
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
di: Choi, Dasol, et al.
Pubblicazione: (2026) -
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
di: Kim, Kihyun, et al.
Pubblicazione: (2026) -
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
di: Lee, DongGeon, et al.
Pubblicazione: (2025)