Residual Tokens Enhance Masked Autoencoders for Speech Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sadok, Samir, Lathuilière, Stéphane, Alameda-Pineda, Xavier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
von: Sadok, Samir, et al.
Veröffentlicht: (2025)
von: Sadok, Samir, et al.
Veröffentlicht: (2025)
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
von: Sadok, Samir, et al.
Veröffentlicht: (2026)
von: Sadok, Samir, et al.
Veröffentlicht: (2026)
Scaling Speech Tokenizers with Diffusion Autoencoders
von: Wang, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2026)
Diffusion-based Unsupervised Audio-visual Speech Enhancement
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2024)
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2024)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Diffusion-based Frameworks for Unsupervised Speech Enhancement
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2026)
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2026)
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
von: Han, Yichen, et al.
Veröffentlicht: (2025)
von: Han, Yichen, et al.
Veröffentlicht: (2025)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification
von: Makineni, Aditya, et al.
Veröffentlicht: (2025)
von: Makineni, Aditya, et al.
Veröffentlicht: (2025)
GSRM: Generative Speech Reward Model for Speech RLHF
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
von: Quelennec, Aurian, et al.
Veröffentlicht: (2025)
von: Quelennec, Aurian, et al.
Veröffentlicht: (2025)
SLM-SS: Speech Language Model for Generative Speech Separation
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
von: Sadeghi, Mostafa, et al.
Veröffentlicht: (2025)
von: Sadeghi, Mostafa, et al.
Veröffentlicht: (2025)
SAME: A Semantically-Aligned Music Autoencoder
von: Parker, Julian D., et al.
Veröffentlicht: (2026)
von: Parker, Julian D., et al.
Veröffentlicht: (2026)
Mask2Flow-TSE: Two-Stage Target Speaker Extraction with Masking and Flow Matching
von: Moon, Junwon, et al.
Veröffentlicht: (2026)
von: Moon, Junwon, et al.
Veröffentlicht: (2026)
Generalizable Speech Deepfake Detection via Information Bottleneck Enhanced Adversarial Alignment
von: Huang, Pu, et al.
Veröffentlicht: (2025)
von: Huang, Pu, et al.
Veröffentlicht: (2025)
Large Speech Model Enabled Semantic Communication
von: Tian, Yun, et al.
Veröffentlicht: (2025)
von: Tian, Yun, et al.
Veröffentlicht: (2025)
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
von: Pappa, Massimiliano, et al.
Veröffentlicht: (2026)
von: Pappa, Massimiliano, et al.
Veröffentlicht: (2026)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
von: Xue, Jun, et al.
Veröffentlicht: (2026)
von: Xue, Jun, et al.
Veröffentlicht: (2026)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification
von: Wu, Bin, et al.
Veröffentlicht: (2024)
von: Wu, Bin, et al.
Veröffentlicht: (2024)
DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
von: Ghosh, Suhita, et al.
Veröffentlicht: (2026)
von: Ghosh, Suhita, et al.
Veröffentlicht: (2026)
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
von: Kim, Kyungsu, et al.
Veröffentlicht: (2025)
von: Kim, Kyungsu, et al.
Veröffentlicht: (2025)
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
von: Kotoge, Rikuto, et al.
Veröffentlicht: (2025)
von: Kotoge, Rikuto, et al.
Veröffentlicht: (2025)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
von: Liu, Jiafeng, et al.
Veröffentlicht: (2026)
von: Liu, Jiafeng, et al.
Veröffentlicht: (2026)
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
von: Rakotoarivony, Lucas
Veröffentlicht: (2026)
von: Rakotoarivony, Lucas
Veröffentlicht: (2026)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
von: Monjur, Mahathir, et al.
Veröffentlicht: (2025)
von: Monjur, Mahathir, et al.
Veröffentlicht: (2025)
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
von: Sadok, Samir, et al.
Veröffentlicht: (2025) -
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
von: Sadok, Samir, et al.
Veröffentlicht: (2026) -
Scaling Speech Tokenizers with Diffusion Autoencoders
von: Wang, Yuancheng, et al.
Veröffentlicht: (2026) -
Diffusion-based Unsupervised Audio-visual Speech Enhancement
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2024) -
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
von: Sadok, Samir, et al.
Veröffentlicht: (2023)