Saved in:
| Main Authors: | Wang, Yichen, Martin, Charles Patrick |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.25207 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
by: Martin, Charles Patrick
Published: (2026)
by: Martin, Charles Patrick
Published: (2026)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
by: Cao, Junjie, et al.
Published: (2025)
by: Cao, Junjie, et al.
Published: (2025)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
Advancing Continual Learning for Robust Deepfake Audio Classification
by: Dong, Feiyi, et al.
Published: (2024)
by: Dong, Feiyi, et al.
Published: (2024)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
by: Zhao, Junqi, et al.
Published: (2025)
by: Zhao, Junqi, et al.
Published: (2025)
Accelerated Convolutive Transfer Function-Based Multichannel NMF Using Iterative Source Steering
by: Xie, Xuemai, et al.
Published: (2025)
by: Xie, Xuemai, et al.
Published: (2025)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
by: Li, Xueyan, et al.
Published: (2025)
by: Li, Xueyan, et al.
Published: (2025)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
by: Ai, Zhiqi, et al.
Published: (2025)
by: Ai, Zhiqi, et al.
Published: (2025)
Differentiable Attenuation Filters for Feedback Delay Networks
by: Ibnyahya, Ilias, et al.
Published: (2025)
by: Ibnyahya, Ilias, et al.
Published: (2025)
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
by: Li, Xiquan, et al.
Published: (2026)
by: Li, Xiquan, et al.
Published: (2026)
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
by: Shen, Huan, et al.
Published: (2026)
by: Shen, Huan, et al.
Published: (2026)
DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation
by: Paar, Ferdinand, et al.
Published: (2026)
by: Paar, Ferdinand, et al.
Published: (2026)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
by: Niu, Xinlei, et al.
Published: (2025)
by: Niu, Xinlei, et al.
Published: (2025)
CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
by: Wang, Siyi, et al.
Published: (2026)
by: Wang, Siyi, et al.
Published: (2026)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
by: Luo, Longjie, et al.
Published: (2025)
by: Luo, Longjie, et al.
Published: (2025)
Dual-Strategy-Enhanced ConBiMamba for Neural Speaker Diarization
by: Liao, Zhen, et al.
Published: (2026)
by: Liao, Zhen, et al.
Published: (2026)
DualMark: Identifying Model and Training Data Origins in Generated Audio
by: Yang, Xuefeng, et al.
Published: (2025)
by: Yang, Xuefeng, et al.
Published: (2025)
CoHear: Conversation Enhancement via Multi-Earphone Collaboration
by: He, Lixing, et al.
Published: (2025)
by: He, Lixing, et al.
Published: (2025)
DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
by: Rika, Daniel, et al.
Published: (2025)
by: Rika, Daniel, et al.
Published: (2025)
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
by: Lee, Yongjoon, et al.
Published: (2026)
by: Lee, Yongjoon, et al.
Published: (2026)
The Shape of Surprise: Structured Uncertainty and Co-Creativity in AI Music Tools
by: Browne, Eric
Published: (2025)
by: Browne, Eric
Published: (2025)
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
by: Dang, Trung, et al.
Published: (2026)
by: Dang, Trung, et al.
Published: (2026)
A Dual-Stage Time-Context Network for Speech-Based Alzheimer's Disease Detection
by: Gao, Yifan, et al.
Published: (2025)
by: Gao, Yifan, et al.
Published: (2025)
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
by: Goto, Keita, et al.
Published: (2026)
by: Goto, Keita, et al.
Published: (2026)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
by: Arora, Siddhant, et al.
Published: (2026)
by: Arora, Siddhant, et al.
Published: (2026)
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
by: Yang, Cheng-Yeh, et al.
Published: (2025)
by: Yang, Cheng-Yeh, et al.
Published: (2025)
MBCodec:Thorough disentangle for high-fidelity audio compression
by: Zhang, Ruonan, et al.
Published: (2025)
by: Zhang, Ruonan, et al.
Published: (2025)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
by: Xing, Jingyuan, et al.
Published: (2025)
by: Xing, Jingyuan, et al.
Published: (2025)
Semantic Audio-Visual Navigation in Continuous Environments
by: Zeng, Yichen, et al.
Published: (2026)
by: Zeng, Yichen, et al.
Published: (2026)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
by: Wang, Yuanyuan, et al.
Published: (2025)
by: Wang, Yuanyuan, et al.
Published: (2025)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Neural Speech Extraction with Human Feedback
by: Itani, Malek, et al.
Published: (2025)
by: Itani, Malek, et al.
Published: (2025)
Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling
by: Negroni, Viola, et al.
Published: (2026)
by: Negroni, Viola, et al.
Published: (2026)
Robust Online Overdetermined Independent Vector Analysis Based on Bilinear Decomposition
by: Chen, Kang, et al.
Published: (2026)
by: Chen, Kang, et al.
Published: (2026)
A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
by: Wu, Yuan-Kuei, et al.
Published: (2025)
by: Wu, Yuan-Kuei, et al.
Published: (2025)
Similar Items
-
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
by: Martin, Charles Patrick
Published: (2026) -
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
by: Niu, Xinlei, et al.
Published: (2024) -
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
by: Cao, Junjie, et al.
Published: (2025) -
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
by: Niu, Xinlei, et al.
Published: (2024) -
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
by: Niu, Xinlei, et al.
Published: (2024)