Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Wanqi, Li, Yanda, Fang, Meng, Wei, Yunchao, Chen, Ling
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912416662552576
author Yang, Wanqi
Li, Yanda
Fang, Meng
Wei, Yunchao
Chen, Ling
author_facet Yang, Wanqi
Li, Yanda
Fang, Meng
Wei, Yunchao
Chen, Ling
contents Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research focused on model-specific adversarial methods, real-world applications demand a more generalizable and universal approach to audio adversarial attacks. In this paper, we introduce the Chat-Audio Attacks (CAA) benchmark including four distinct types of audio attacks, which aims to explore the vulnerabilities of LALMs to these audio attacks in conversational scenarios. To evaluate the robustness of LALMs, we propose three evaluation strategies: Standard Evaluation, utilizing traditional metrics to quantify model performance under attacks; GPT-4o-Based Evaluation, which simulates real-world conversational complexities; and Human Evaluation, offering insights into user perception and trust. We evaluate six state-of-the-art LALMs with voice interaction capabilities, including Gemini-1.5-Pro, GPT-4o, and others, using three distinct evaluation methods on the CAA benchmark. Our comprehensive analysis reveals the impact of four types of audio attacks on the performance of these models, demonstrating that GPT-4o exhibits the highest level of resilience. Our data can be accessed via the following link: \href{https://github.com/crystraldo/CAA}{CAA}.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
Yang, Wanqi
Li, Yanda
Fang, Meng
Wei, Yunchao
Chen, Ling
Sound
Artificial Intelligence
Audio and Speech Processing
Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research focused on model-specific adversarial methods, real-world applications demand a more generalizable and universal approach to audio adversarial attacks. In this paper, we introduce the Chat-Audio Attacks (CAA) benchmark including four distinct types of audio attacks, which aims to explore the vulnerabilities of LALMs to these audio attacks in conversational scenarios. To evaluate the robustness of LALMs, we propose three evaluation strategies: Standard Evaluation, utilizing traditional metrics to quantify model performance under attacks; GPT-4o-Based Evaluation, which simulates real-world conversational complexities; and Human Evaluation, offering insights into user perception and trust. We evaluate six state-of-the-art LALMs with voice interaction capabilities, including Gemini-1.5-Pro, GPT-4o, and others, using three distinct evaluation methods on the CAA benchmark. Our comprehensive analysis reveals the impact of four types of audio attacks on the performance of these models, demonstrating that GPT-4o exhibits the highest level of resilience. Our data can be accessed via the following link: \href{https://github.com/crystraldo/CAA}{CAA}.
title Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2411.14842