Audio Generation Through Score-Based Generative Modeling: Design Principles and Implementation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Ge, Wen, Yutong, Duan, Zhiyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
A Probabilistic Fusion Framework for Spoofing Aware Speaker Verification
von: Zhang, You, et al.
Veröffentlicht: (2022)
von: Zhang, You, et al.
Veröffentlicht: (2022)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021
von: Chen, Xinhui, et al.
Veröffentlicht: (2021)
von: Chen, Xinhui, et al.
Veröffentlicht: (2021)
An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
von: Zhang, You, et al.
Veröffentlicht: (2021)
von: Zhang, You, et al.
Veröffentlicht: (2021)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
A Generative-First Neural Audio Autoencoder
von: Casebeer, Jonah, et al.
Veröffentlicht: (2026)
von: Casebeer, Jonah, et al.
Veröffentlicht: (2026)
Scoring Time Intervals using Non-Hierarchical Transformer For Automatic Piano Transcription
von: Yan, Yujia, et al.
Veröffentlicht: (2024)
von: Yan, Yujia, et al.
Veröffentlicht: (2024)
PromptSep: Generative Audio Separation via Multimodal Prompting
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Latent Watermarking of Audio Generative Models
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2026)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2026)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
Local Density-Based Anomaly Score Normalization for Domain Generalization
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
von: Zhang, You, et al.
Veröffentlicht: (2025)
von: Zhang, You, et al.
Veröffentlicht: (2025)
Predicting Global HRTFs From Scanned Head Geometry Using Deep Learning and Compact Representations
von: Wang, Yuxiang, et al.
Veröffentlicht: (2022)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2022)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2026)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Probing Audio-Generation Capabilities of Text-Based Language Models
von: Anbazhagan, Arjun Prasaath, et al.
Veröffentlicht: (2025)
von: Anbazhagan, Arjun Prasaath, et al.
Veröffentlicht: (2025)
DAC-JAX: A JAX Implementation of the Descript Audio Codec
von: Braun, David
Veröffentlicht: (2024)
von: Braun, David
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024) -
A Probabilistic Fusion Framework for Spoofing Aware Speaker Verification
von: Zhang, You, et al.
Veröffentlicht: (2022) -
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025) -
UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021
von: Chen, Xinhui, et al.
Veröffentlicht: (2021) -
An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
von: Zhang, You, et al.
Veröffentlicht: (2021)