Gespeichert in:
| Hauptverfasser: | Bahadi, Soufiyan, Plourde, Eric, Rouat, Jean |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.08188 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Central Frequencies Locally Competitive Algorithm for Speech
von: Bahadi, Soufiyan, et al.
Veröffentlicht: (2025)
von: Bahadi, Soufiyan, et al.
Veröffentlicht: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
von: Landau, Gilad, et al.
Veröffentlicht: (2025)
von: Landau, Gilad, et al.
Veröffentlicht: (2025)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
SpatialCodec: Neural Spatial Speech Coding
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
von: Büthe, Jan, et al.
Veröffentlicht: (2023)
von: Büthe, Jan, et al.
Veröffentlicht: (2023)
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
A Multilingual Framework for Dysarthria: Detection, Severity Classification, Speech-to-Text, and Clean Speech Generation
von: Raghu, Ananya, et al.
Veröffentlicht: (2025)
von: Raghu, Ananya, et al.
Veröffentlicht: (2025)
VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Adaptive Convolution for CNN-based Speech Enhancement Models
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
von: Yuan, Xihao, et al.
Veröffentlicht: (2025)
von: Yuan, Xihao, et al.
Veröffentlicht: (2025)
UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension
von: Gupta, Kishan, et al.
Veröffentlicht: (2025)
von: Gupta, Kishan, et al.
Veröffentlicht: (2025)
Fewer-token Neural Speech Codec with Time-invariant Codes
von: Ren, Yong, et al.
Veröffentlicht: (2023)
von: Ren, Yong, et al.
Veröffentlicht: (2023)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2024)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2024)
Adaptive Slimming for Scalable and Efficient Speech Enhancement
von: Miccini, Riccardo, et al.
Veröffentlicht: (2025)
von: Miccini, Riccardo, et al.
Veröffentlicht: (2025)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
Synthetic Speech Classification: IEEE Signal Processing Cup 2022 challenge
von: Rahmun, Mahieyin, et al.
Veröffentlicht: (2024)
von: Rahmun, Mahieyin, et al.
Veröffentlicht: (2024)
Adaptive Differential Denoising for Respiratory Sounds Classification
von: Dong, Gaoyang, et al.
Veröffentlicht: (2025)
von: Dong, Gaoyang, et al.
Veröffentlicht: (2025)
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
von: Li, Sirui, et al.
Veröffentlicht: (2025)
von: Li, Sirui, et al.
Veröffentlicht: (2025)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
Combined Generative and Predictive Modeling for Speech Super-resolution
von: Wang, Heming, et al.
Veröffentlicht: (2024)
von: Wang, Heming, et al.
Veröffentlicht: (2024)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
von: Wan, Zixiang, et al.
Veröffentlicht: (2025)
von: Wan, Zixiang, et al.
Veröffentlicht: (2025)
On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
von: Hao, Yaqian, et al.
Veröffentlicht: (2024)
von: Hao, Yaqian, et al.
Veröffentlicht: (2024)
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2025)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2025)
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adaptive Central Frequencies Locally Competitive Algorithm for Speech
von: Bahadi, Soufiyan, et al.
Veröffentlicht: (2025) -
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024) -
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024) -
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025) -
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)