SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chih-Kai, Piao, Yen-Ting, Hsu, Tzu-Wen, Fu, Szu-Wei, Chen, Zhehuai, Lu, Ke-Han, Huang, Sung-Feng, Yang, Chao-Han Huck, Wang, Yu-Chiang Frank, Chen, Yun-Nung, Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915863669506048
author Yang, Chih-Kai
Piao, Yen-Ting
Hsu, Tzu-Wen
Fu, Szu-Wei
Chen, Zhehuai
Lu, Ke-Han
Huang, Sung-Feng
Yang, Chao-Han Huck
Wang, Yu-Chiang Frank
Chen, Yun-Nung
Lee, Hung-yi
author_facet Yang, Chih-Kai
Piao, Yen-Ting
Hsu, Tzu-Wen
Fu, Szu-Wei
Chen, Zhehuai
Lu, Ke-Han
Huang, Sung-Feng
Yang, Chao-Han Huck
Wang, Yu-Chiang Frank
Chen, Yun-Nung
Lee, Hung-yi
contents Knowledge editing enables targeted updates without retraining, but prior work focuses on textual or visual facts, leaving abstract auditory perceptual knowledge underexplored. We introduce SAKE, the first benchmark for editing perceptual auditory attribute knowledge in large audio-language models (LALMs), which requires modifying acoustic generalization rather than isolated facts. We evaluate eight diverse editing methods on three LALMs across reliability, generality, locality, and portability, under single and sequential edits. Results show that most methods enforce edits reliably but struggle with auditory generalization, intra-attribute locality, and multimodal knowledge propagation, and often exhibit forgetting or degeneration in sequential editing. Additionally, fine-tuning the modality connector emerges as a more robust and balanced baseline compared with directly editing the LLM backbones. SAKE reveals key limitations of current methods and provides a foundation for developing auditory-specific LALM editing techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16917
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
Yang, Chih-Kai
Piao, Yen-Ting
Hsu, Tzu-Wen
Fu, Szu-Wei
Chen, Zhehuai
Lu, Ke-Han
Huang, Sung-Feng
Yang, Chao-Han Huck
Wang, Yu-Chiang Frank
Chen, Yun-Nung
Lee, Hung-yi
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Knowledge editing enables targeted updates without retraining, but prior work focuses on textual or visual facts, leaving abstract auditory perceptual knowledge underexplored. We introduce SAKE, the first benchmark for editing perceptual auditory attribute knowledge in large audio-language models (LALMs), which requires modifying acoustic generalization rather than isolated facts. We evaluate eight diverse editing methods on three LALMs across reliability, generality, locality, and portability, under single and sequential edits. Results show that most methods enforce edits reliably but struggle with auditory generalization, intra-attribute locality, and multimodal knowledge propagation, and often exhibit forgetting or degeneration in sequential editing. Additionally, fine-tuning the modality connector emerges as a more robust and balanced baseline compared with directly editing the LLM backbones. SAKE reveals key limitations of current methods and provides a foundation for developing auditory-specific LALM editing techniques.
title SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2510.16917