AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ren, Yiming, Xu, Xuenan, Zhang, Ziyang, Wu, Wen, Li, Baoxiang, Zhang, Chao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918512801349632
author Ren, Yiming
Xu, Xuenan
Zhang, Ziyang
Wu, Wen
Li, Baoxiang
Zhang, Chao
author_facet Ren, Yiming
Xu, Xuenan
Zhang, Ziyang
Wu, Wen
Li, Baoxiang
Zhang, Chao
contents Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched character settings with voice performance, insufficient self-correction mechanisms, and limited human interactivity. To address these challenges, we propose AuDirector, a self-reflective closed-loop multi-agent framework. Specifically, it involves an Identity-Aware Pre-production mechanism that transforms narrative texts into character profiles and utterance-level emotional instructions to retrieve suitable voice candidates and guide expressive speech synthesis, thereby promoting context-aligned voice adaptation. To enhance quality, a Collaborative Synthesis and Correction module introduces a closed-loop self-correction mechanism to systematically audit and regenerate defective audio components. Furthermore, a Human-Guided Interactive Refinement module facilitates user control by interpreting natural language feedback to interactively refine the underlying scripts. Experiments demonstrate that AuDirector achieves superior performance compared to state-of-the-art baselines in structural coherence, emotional expressiveness, and acoustic fidelity. Audio samples can be found at https://anonymous-itsh.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11866
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
Ren, Yiming
Xu, Xuenan
Zhang, Ziyang
Wu, Wen
Li, Baoxiang
Zhang, Chao
Sound
Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched character settings with voice performance, insufficient self-correction mechanisms, and limited human interactivity. To address these challenges, we propose AuDirector, a self-reflective closed-loop multi-agent framework. Specifically, it involves an Identity-Aware Pre-production mechanism that transforms narrative texts into character profiles and utterance-level emotional instructions to retrieve suitable voice candidates and guide expressive speech synthesis, thereby promoting context-aligned voice adaptation. To enhance quality, a Collaborative Synthesis and Correction module introduces a closed-loop self-correction mechanism to systematically audit and regenerate defective audio components. Furthermore, a Human-Guided Interactive Refinement module facilitates user control by interpreting natural language feedback to interactively refine the underlying scripts. Experiments demonstrate that AuDirector achieves superior performance compared to state-of-the-art baselines in structural coherence, emotional expressiveness, and acoustic fidelity. Audio samples can be found at https://anonymous-itsh.github.io/.
title AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
topic Sound
url https://arxiv.org/abs/2605.11866