One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Stylianou, Ioannis, Francombe, Jon, Martinez-Nuevo, Pablo, Shepstone, Sven Ewan, Tan, Zheng-Hua
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917482562846720
author Stylianou, Ioannis
Francombe, Jon
Martinez-Nuevo, Pablo
Shepstone, Sven Ewan
Tan, Zheng-Hua
author_facet Stylianou, Ioannis
Francombe, Jon
Martinez-Nuevo, Pablo
Shepstone, Sven Ewan
Tan, Zheng-Hua
contents Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts (e.g., mood, location, or social setting). In this paper, we introduce a Large Language Model (LLM)-based alternative that maps natural language text prompts to equalization settings. This enables a conversational approach to sound system control. By utilizing data collected from a controlled listening experiment, our models exploit in-context learning and parameter-efficient fine-tuning techniques to reliably align with population-preferred equalization settings. Our evaluation methods, which leverage distributional metrics that capture users' varied preferences, show statistically significant improvements in distributional alignment over random sampling and static preset baselines. These results indicate that LLMs could function as "artificial equalizers," contributing to the development of more accessible, context-aware, and expert-level audio tuning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_09448
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
Stylianou, Ioannis
Francombe, Jon
Martinez-Nuevo, Pablo
Shepstone, Sven Ewan
Tan, Zheng-Hua
Sound
Artificial Intelligence
Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts (e.g., mood, location, or social setting). In this paper, we introduce a Large Language Model (LLM)-based alternative that maps natural language text prompts to equalization settings. This enables a conversational approach to sound system control. By utilizing data collected from a controlled listening experiment, our models exploit in-context learning and parameter-efficient fine-tuning techniques to reliably align with population-preferred equalization settings. Our evaluation methods, which leverage distributional metrics that capture users' varied preferences, show statistically significant improvements in distributional alignment over random sampling and static preset baselines. These results indicate that LLMs could function as "artificial equalizers," contributing to the development of more accessible, context-aware, and expert-level audio tuning methods.
title One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2601.09448