Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Minje, Skoglund, Jan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
von: Kuznetsova, Anastasia, et al.
Veröffentlicht: (2025)
von: Kuznetsova, Anastasia, et al.
Veröffentlicht: (2025)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
Latent Granular Resynthesis using Neural Audio Codecs
von: Tokui, Nao, et al.
Veröffentlicht: (2025)
von: Tokui, Nao, et al.
Veröffentlicht: (2025)
Speech Enhancement Based on Drifting Models
von: Xu, Liang, et al.
Veröffentlicht: (2026)
von: Xu, Liang, et al.
Veröffentlicht: (2026)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
SpectroStream: A Versatile Neural Codec for General Audio
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
AI-Generated Music Detection in Broadcast Monitoring
von: López-Ayala, David, et al.
Veröffentlicht: (2026)
von: López-Ayala, David, et al.
Veröffentlicht: (2026)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
On the Design of Diffusion-based Neural Speech Codecs
von: Foti, Pietro, et al.
Veröffentlicht: (2025)
von: Foti, Pietro, et al.
Veröffentlicht: (2025)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
von: Steckel, Jan, et al.
Veröffentlicht: (2024)
von: Steckel, Jan, et al.
Veröffentlicht: (2024)
Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
von: Wilroth, Johanna, et al.
Veröffentlicht: (2026)
von: Wilroth, Johanna, et al.
Veröffentlicht: (2026)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Aliasing-Free Neural Audio Synthesis
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
Recent Advances in Discrete Speech Tokens: A Review
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024) -
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024) -
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024) -
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
von: Kuznetsova, Anastasia, et al.
Veröffentlicht: (2025) -
Compressing Quaternion Convolutional Neural Networks for Audio Classification
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)