Huí Sù: Co-constructing a Dual Feedback Apparatus
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yichen, Martin, Charles Patrick |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
di: Martin, Charles Patrick
Pubblicazione: (2026)
di: Martin, Charles Patrick
Pubblicazione: (2026)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
di: Cao, Junjie, et al.
Pubblicazione: (2025)
di: Cao, Junjie, et al.
Pubblicazione: (2025)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
di: Zhao, Junqi, et al.
Pubblicazione: (2025)
di: Zhao, Junqi, et al.
Pubblicazione: (2025)
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
di: Ai, Zhiqi, et al.
Pubblicazione: (2025)
di: Ai, Zhiqi, et al.
Pubblicazione: (2025)
Advancing Continual Learning for Robust Deepfake Audio Classification
di: Dong, Feiyi, et al.
Pubblicazione: (2024)
di: Dong, Feiyi, et al.
Pubblicazione: (2024)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
di: Li, Xueyan, et al.
Pubblicazione: (2025)
di: Li, Xueyan, et al.
Pubblicazione: (2025)
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
di: Li, Xiquan, et al.
Pubblicazione: (2026)
di: Li, Xiquan, et al.
Pubblicazione: (2026)
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
di: Shen, Huan, et al.
Pubblicazione: (2026)
di: Shen, Huan, et al.
Pubblicazione: (2026)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
di: Wang, Zehan, et al.
Pubblicazione: (2025)
di: Wang, Zehan, et al.
Pubblicazione: (2025)
Differentiable Attenuation Filters for Feedback Delay Networks
di: Ibnyahya, Ilias, et al.
Pubblicazione: (2025)
di: Ibnyahya, Ilias, et al.
Pubblicazione: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
di: Li, Tao, et al.
Pubblicazione: (2025)
di: Li, Tao, et al.
Pubblicazione: (2025)
Accelerated Convolutive Transfer Function-Based Multichannel NMF Using Iterative Source Steering
di: Xie, Xuemai, et al.
Pubblicazione: (2025)
di: Xie, Xuemai, et al.
Pubblicazione: (2025)
Dual-Strategy-Enhanced ConBiMamba for Neural Speaker Diarization
di: Liao, Zhen, et al.
Pubblicazione: (2026)
di: Liao, Zhen, et al.
Pubblicazione: (2026)
DualMark: Identifying Model and Training Data Origins in Generated Audio
di: Yang, Xuefeng, et al.
Pubblicazione: (2025)
di: Yang, Xuefeng, et al.
Pubblicazione: (2025)
CoHear: Conversation Enhancement via Multi-Earphone Collaboration
di: He, Lixing, et al.
Pubblicazione: (2025)
di: He, Lixing, et al.
Pubblicazione: (2025)
DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
di: Rika, Daniel, et al.
Pubblicazione: (2025)
di: Rika, Daniel, et al.
Pubblicazione: (2025)
DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation
di: Paar, Ferdinand, et al.
Pubblicazione: (2026)
di: Paar, Ferdinand, et al.
Pubblicazione: (2026)
CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
di: Wang, Siyi, et al.
Pubblicazione: (2026)
di: Wang, Siyi, et al.
Pubblicazione: (2026)
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
di: Qiu, Jielin, et al.
Pubblicazione: (2026)
di: Qiu, Jielin, et al.
Pubblicazione: (2026)
The Shape of Surprise: Structured Uncertainty and Co-Creativity in AI Music Tools
di: Browne, Eric
Pubblicazione: (2025)
di: Browne, Eric
Pubblicazione: (2025)
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
di: Dang, Trung, et al.
Pubblicazione: (2026)
di: Dang, Trung, et al.
Pubblicazione: (2026)
A Dual-Stage Time-Context Network for Speech-Based Alzheimer's Disease Detection
di: Gao, Yifan, et al.
Pubblicazione: (2025)
di: Gao, Yifan, et al.
Pubblicazione: (2025)
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
di: Goto, Keita, et al.
Pubblicazione: (2026)
di: Goto, Keita, et al.
Pubblicazione: (2026)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2025)
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2025)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
di: Arora, Siddhant, et al.
Pubblicazione: (2026)
di: Arora, Siddhant, et al.
Pubblicazione: (2026)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
di: Luo, Longjie, et al.
Pubblicazione: (2025)
di: Luo, Longjie, et al.
Pubblicazione: (2025)
Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling
di: Negroni, Viola, et al.
Pubblicazione: (2026)
di: Negroni, Viola, et al.
Pubblicazione: (2026)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling
di: Lin, Rui, et al.
Pubblicazione: (2025)
di: Lin, Rui, et al.
Pubblicazione: (2025)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
di: Li, Zongyi, et al.
Pubblicazione: (2025)
di: Li, Zongyi, et al.
Pubblicazione: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
di: Xie, Hanke, et al.
Pubblicazione: (2025)
di: Xie, Hanke, et al.
Pubblicazione: (2025)
A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
di: Wu, Yuan-Kuei, et al.
Pubblicazione: (2025)
di: Wu, Yuan-Kuei, et al.
Pubblicazione: (2025)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
di: Wang, Haoxu, et al.
Pubblicazione: (2025)
di: Wang, Haoxu, et al.
Pubblicazione: (2025)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
di: Martin, Charles Patrick
Pubblicazione: (2026) -
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
di: Niu, Xinlei, et al.
Pubblicazione: (2024) -
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
di: Cao, Junjie, et al.
Pubblicazione: (2025) -
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
di: Niu, Xinlei, et al.
Pubblicazione: (2024) -
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
di: Zhao, Junqi, et al.
Pubblicazione: (2025)