MaskSR: Masked Language Model for Full-band Speech Restoration
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xu, Wang, Qirui, Liu, Xiaoyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
di: Wang, Yuancheng, et al.
Pubblicazione: (2025)
di: Wang, Yuancheng, et al.
Pubblicazione: (2025)
Deep Active Speech Cancellation with Mamba-Masking Network
di: Mishaly, Yehuda, et al.
Pubblicazione: (2025)
di: Mishaly, Yehuda, et al.
Pubblicazione: (2025)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
di: Parker, Julian D, et al.
Pubblicazione: (2024)
di: Parker, Julian D, et al.
Pubblicazione: (2024)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
di: Ziogas, Ioannis, et al.
Pubblicazione: (2024)
di: Ziogas, Ioannis, et al.
Pubblicazione: (2024)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
di: Sinha, Abhijit, et al.
Pubblicazione: (2025)
di: Sinha, Abhijit, et al.
Pubblicazione: (2025)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
di: Baoueb, Teysir, et al.
Pubblicazione: (2025)
di: Baoueb, Teysir, et al.
Pubblicazione: (2025)
Speech Enhancement Based on Drifting Models
di: Xu, Liang, et al.
Pubblicazione: (2026)
di: Xu, Liang, et al.
Pubblicazione: (2026)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
di: Yao, Shengshi, et al.
Pubblicazione: (2025)
di: Yao, Shengshi, et al.
Pubblicazione: (2025)
Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
di: Drossos, Konstantinos, et al.
Pubblicazione: (2025)
di: Drossos, Konstantinos, et al.
Pubblicazione: (2025)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
di: Choi, Woongjib, et al.
Pubblicazione: (2025)
di: Choi, Woongjib, et al.
Pubblicazione: (2025)
An Explainable Proxy Model for Multiabel Audio Segmentation
di: Mariotte, Théo, et al.
Pubblicazione: (2024)
di: Mariotte, Théo, et al.
Pubblicazione: (2024)
Model as Loss: A Self-Consistent Training Paradigm
di: Phaye, Saisamarth Rajesh, et al.
Pubblicazione: (2025)
di: Phaye, Saisamarth Rajesh, et al.
Pubblicazione: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
di: Chung, Yoonjin, et al.
Pubblicazione: (2024)
di: Chung, Yoonjin, et al.
Pubblicazione: (2024)
TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
di: Huang, Zhaolan, et al.
Pubblicazione: (2024)
di: Huang, Zhaolan, et al.
Pubblicazione: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
di: Wu, Linzhi, et al.
Pubblicazione: (2026)
di: Wu, Linzhi, et al.
Pubblicazione: (2026)
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
di: Baoueb, Teysir, et al.
Pubblicazione: (2025)
di: Baoueb, Teysir, et al.
Pubblicazione: (2025)
High-Resolution Speech Restoration with Latent Diffusion Model
di: Dhyani, Tushar, et al.
Pubblicazione: (2024)
di: Dhyani, Tushar, et al.
Pubblicazione: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
di: Ting, Zhu, et al.
Pubblicazione: (2024)
di: Ting, Zhu, et al.
Pubblicazione: (2024)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
di: Mancini, Eleonora, et al.
Pubblicazione: (2024)
di: Mancini, Eleonora, et al.
Pubblicazione: (2024)
Gull: A Generative Multifunctional Audio Codec
di: Luo, Yi, et al.
Pubblicazione: (2024)
di: Luo, Yi, et al.
Pubblicazione: (2024)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
di: Kim, Hounsu, et al.
Pubblicazione: (2024)
di: Kim, Hounsu, et al.
Pubblicazione: (2024)
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
di: Kim, Daewoong, et al.
Pubblicazione: (2024)
di: Kim, Daewoong, et al.
Pubblicazione: (2024)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
di: Ali, Shams Nafisa, et al.
Pubblicazione: (2024)
di: Ali, Shams Nafisa, et al.
Pubblicazione: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
Real-time Timbre Remapping with Differentiable DSP
di: Shier, Jordie, et al.
Pubblicazione: (2024)
di: Shier, Jordie, et al.
Pubblicazione: (2024)
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
di: Bitra, Venkat Suprabath, et al.
Pubblicazione: (2026)
di: Bitra, Venkat Suprabath, et al.
Pubblicazione: (2026)
Design Of Rubble Analyzer Probe Using ML For Earthquake
di: Sebastian, Abhishek, et al.
Pubblicazione: (2023)
di: Sebastian, Abhishek, et al.
Pubblicazione: (2023)
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
di: Ackva, Valentin, et al.
Pubblicazione: (2025)
di: Ackva, Valentin, et al.
Pubblicazione: (2025)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
di: Tuncay, Ludovic, et al.
Pubblicazione: (2025)
di: Tuncay, Ludovic, et al.
Pubblicazione: (2025)
Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
di: Di Carlo, Diego, et al.
Pubblicazione: (2025)
di: Di Carlo, Diego, et al.
Pubblicazione: (2025)
Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
di: Limberg, Christian, et al.
Pubblicazione: (2025)
di: Limberg, Christian, et al.
Pubblicazione: (2025)
AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
di: Siddiqui, Md. Saiful Bari, et al.
Pubblicazione: (2025)
di: Siddiqui, Md. Saiful Bari, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024) -
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
di: Wang, Yuancheng, et al.
Pubblicazione: (2025) -
Deep Active Speech Cancellation with Mamba-Masking Network
di: Mishaly, Yehuda, et al.
Pubblicazione: (2025) -
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
di: Gállego, Gerard I., et al.
Pubblicazione: (2024) -
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)