On the Design of Diffusion-based Neural Speech Codecs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Foti, Pietro, Brendel, Andreas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
Analyzing the Impact of Splicing Artifacts in Partially Fake Speech Signals
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Neural Style Transfer for Audio Spectograms
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features
von: Luo, Jing, et al.
Veröffentlicht: (2024)
von: Luo, Jing, et al.
Veröffentlicht: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Recent Advances in Discrete Speech Tokens: A Review
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
CoComposer: LLM Multi-agent Collaborative Music Composition
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
von: Seo, Jinbae, et al.
Veröffentlicht: (2025)
von: Seo, Jinbae, et al.
Veröffentlicht: (2025)
Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings
von: Shah, Ankit, et al.
Veröffentlicht: (2025)
von: Shah, Ankit, et al.
Veröffentlicht: (2025)
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
von: Wu, Junxian, et al.
Veröffentlicht: (2025)
von: Wu, Junxian, et al.
Veröffentlicht: (2025)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
Embedding Alignment in Code Generation for Audio
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
From Sound to Sight: Towards AI-authored Music Videos
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
YuE: Scaling Open Foundation Models for Long-Form Music Generation
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
Prompt-aware classifier free guidance for diffusion models
von: Zhang, Xuanhao, et al.
Veröffentlicht: (2025)
von: Zhang, Xuanhao, et al.
Veröffentlicht: (2025)
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
von: Zeng, Wei, et al.
Veröffentlicht: (2025)
von: Zeng, Wei, et al.
Veröffentlicht: (2025)
Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
von: Luo, Jing, et al.
Veröffentlicht: (2025)
von: Luo, Jing, et al.
Veröffentlicht: (2025)
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
von: Klemt, Marcel, et al.
Veröffentlicht: (2025)
von: Klemt, Marcel, et al.
Veröffentlicht: (2025)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
von: Li, Maomao, et al.
Veröffentlicht: (2026)
von: Li, Maomao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024) -
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023) -
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
von: Lou, Haowei, et al.
Veröffentlicht: (2024) -
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025) -
Analyzing the Impact of Splicing Artifacts in Partially Fake Speech Signals
von: Negroni, Viola, et al.
Veröffentlicht: (2024)