Huí Sù: Co-constructing a Dual Feedback Apparatus
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yichen, Martin, Charles Patrick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
von: Martin, Charles Patrick
Veröffentlicht: (2026)
von: Martin, Charles Patrick
Veröffentlicht: (2026)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
von: Ai, Zhiqi, et al.
Veröffentlicht: (2025)
von: Ai, Zhiqi, et al.
Veröffentlicht: (2025)
Advancing Continual Learning for Robust Deepfake Audio Classification
von: Dong, Feiyi, et al.
Veröffentlicht: (2024)
von: Dong, Feiyi, et al.
Veröffentlicht: (2024)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
von: Li, Xueyan, et al.
Veröffentlicht: (2025)
von: Li, Xueyan, et al.
Veröffentlicht: (2025)
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
von: Li, Xiquan, et al.
Veröffentlicht: (2026)
von: Li, Xiquan, et al.
Veröffentlicht: (2026)
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
von: Shen, Huan, et al.
Veröffentlicht: (2026)
von: Shen, Huan, et al.
Veröffentlicht: (2026)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
Differentiable Attenuation Filters for Feedback Delay Networks
von: Ibnyahya, Ilias, et al.
Veröffentlicht: (2025)
von: Ibnyahya, Ilias, et al.
Veröffentlicht: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
Accelerated Convolutive Transfer Function-Based Multichannel NMF Using Iterative Source Steering
von: Xie, Xuemai, et al.
Veröffentlicht: (2025)
von: Xie, Xuemai, et al.
Veröffentlicht: (2025)
Dual-Strategy-Enhanced ConBiMamba for Neural Speaker Diarization
von: Liao, Zhen, et al.
Veröffentlicht: (2026)
von: Liao, Zhen, et al.
Veröffentlicht: (2026)
DualMark: Identifying Model and Training Data Origins in Generated Audio
von: Yang, Xuefeng, et al.
Veröffentlicht: (2025)
von: Yang, Xuefeng, et al.
Veröffentlicht: (2025)
CoHear: Conversation Enhancement via Multi-Earphone Collaboration
von: He, Lixing, et al.
Veröffentlicht: (2025)
von: He, Lixing, et al.
Veröffentlicht: (2025)
DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
von: Rika, Daniel, et al.
Veröffentlicht: (2025)
von: Rika, Daniel, et al.
Veröffentlicht: (2025)
DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation
von: Paar, Ferdinand, et al.
Veröffentlicht: (2026)
von: Paar, Ferdinand, et al.
Veröffentlicht: (2026)
CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
von: Wang, Siyi, et al.
Veröffentlicht: (2026)
von: Wang, Siyi, et al.
Veröffentlicht: (2026)
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
von: Qiu, Jielin, et al.
Veröffentlicht: (2026)
von: Qiu, Jielin, et al.
Veröffentlicht: (2026)
The Shape of Surprise: Structured Uncertainty and Co-Creativity in AI Music Tools
von: Browne, Eric
Veröffentlicht: (2025)
von: Browne, Eric
Veröffentlicht: (2025)
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
von: Dang, Trung, et al.
Veröffentlicht: (2026)
von: Dang, Trung, et al.
Veröffentlicht: (2026)
A Dual-Stage Time-Context Network for Speech-Based Alzheimer's Disease Detection
von: Gao, Yifan, et al.
Veröffentlicht: (2025)
von: Gao, Yifan, et al.
Veröffentlicht: (2025)
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
von: Goto, Keita, et al.
Veröffentlicht: (2026)
von: Goto, Keita, et al.
Veröffentlicht: (2026)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
von: Lee, Yongjoon, et al.
Veröffentlicht: (2026)
von: Lee, Yongjoon, et al.
Veröffentlicht: (2026)
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2025)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2025)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
von: Arora, Siddhant, et al.
Veröffentlicht: (2026)
von: Arora, Siddhant, et al.
Veröffentlicht: (2026)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling
von: Negroni, Viola, et al.
Veröffentlicht: (2026)
von: Negroni, Viola, et al.
Veröffentlicht: (2026)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling
von: Lin, Rui, et al.
Veröffentlicht: (2025)
von: Lin, Rui, et al.
Veröffentlicht: (2025)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
von: Wu, Yuan-Kuei, et al.
Veröffentlicht: (2025)
von: Wu, Yuan-Kuei, et al.
Veröffentlicht: (2025)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
von: Martin, Charles Patrick
Veröffentlicht: (2026) -
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024) -
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025) -
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
von: Niu, Xinlei, et al.
Veröffentlicht: (2024) -
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)