Efficient Parallel Audio Generation using Group Masked Language Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeong, Myeonghun, Kim, Minchan, Lee, Joun Yeop, Kim, Nam Soo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024)
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
von: Kim, Semin, et al.
Veröffentlicht: (2024)
von: Kim, Semin, et al.
Veröffentlicht: (2024)
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
von: Chung, Yoonjin, et al.
Veröffentlicht: (2025)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2025)
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
AudioGenX: Explainability on Text-to-Audio Generative Models
von: Kang, Hyunju, et al.
Veröffentlicht: (2025)
von: Kang, Hyunju, et al.
Veröffentlicht: (2025)
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
Guiding Audio Editing with Audio Language Model
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription
von: Kim, Hounsu, et al.
Veröffentlicht: (2025)
von: Kim, Hounsu, et al.
Veröffentlicht: (2025)
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
von: Kang, Ju Yeon, et al.
Veröffentlicht: (2025)
von: Kang, Ju Yeon, et al.
Veröffentlicht: (2025)
Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
von: Jeong, Iksoon, et al.
Veröffentlicht: (2025)
von: Jeong, Iksoon, et al.
Veröffentlicht: (2025)
On the de-duplication of the Lakh MIDI dataset
von: Choi, Eunjin, et al.
Veröffentlicht: (2025)
von: Choi, Eunjin, et al.
Veröffentlicht: (2025)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
von: Lee, Hyeonseung, et al.
Veröffentlicht: (2024)
von: Lee, Hyeonseung, et al.
Veröffentlicht: (2024)
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2026)
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2026)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
von: Long, Phillip, et al.
Veröffentlicht: (2026)
von: Long, Phillip, et al.
Veröffentlicht: (2026)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
von: Luebs, Alejandro, et al.
Veröffentlicht: (2026)
von: Luebs, Alejandro, et al.
Veröffentlicht: (2026)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
von: Cohen, Ohad, et al.
Veröffentlicht: (2024)
von: Cohen, Ohad, et al.
Veröffentlicht: (2024)
LongAudio-RAG: Event-Grounded Question Answering over Multi-Hour Long Audio
von: Vakada, Naveen, et al.
Veröffentlicht: (2026)
von: Vakada, Naveen, et al.
Veröffentlicht: (2026)
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
von: Passoni, Riccardo, et al.
Veröffentlicht: (2025)
von: Passoni, Riccardo, et al.
Veröffentlicht: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
A Survey of Deep Learning Audio Generation Methods
von: Božić, Matej, et al.
Veröffentlicht: (2024)
von: Božić, Matej, et al.
Veröffentlicht: (2024)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
von: Jang, Jaehyuk, et al.
Veröffentlicht: (2026)
von: Jang, Jaehyuk, et al.
Veröffentlicht: (2026)
Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning
von: Luong, Manh, et al.
Veröffentlicht: (2025)
von: Luong, Manh, et al.
Veröffentlicht: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations
von: Grinberg, Petr, et al.
Veröffentlicht: (2025)
von: Grinberg, Petr, et al.
Veröffentlicht: (2025)
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Prompt-guided Precise Audio Editing with Diffusion Models
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
von: Song, Yakun, et al.
Veröffentlicht: (2025)
von: Song, Yakun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
von: Kim, Minchan, et al.
Veröffentlicht: (2024) -
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024) -
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024) -
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
von: Kim, Semin, et al.
Veröffentlicht: (2024) -
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
von: Chung, Yoonjin, et al.
Veröffentlicht: (2025)