Wanna hear your voice? A sample is all we need!
Fuente:
arXiv
Salvato in:
| Autori principali: | Pham, The Hieu, Nguyen, Phuong Thanh Tran, Nguyen, Xuan Tho, Nguyen, Tan Dat, Nguyen, Duc Dung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
PhoWhisper: Automatic Speech Recognition for Vietnamese
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
di: Lam, Phat, et al.
Pubblicazione: (2024)
di: Lam, Phat, et al.
Pubblicazione: (2024)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
di: Nguyen, Tin, et al.
Pubblicazione: (2024)
di: Nguyen, Tin, et al.
Pubblicazione: (2024)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations
di: Trinh, Quoc-Huy, et al.
Pubblicazione: (2024)
di: Trinh, Quoc-Huy, et al.
Pubblicazione: (2024)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2026)
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2026)
Continuous Learning of Transformer-based Audio Deepfake Detection
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
di: Ho, Luong, et al.
Pubblicazione: (2025)
di: Ho, Luong, et al.
Pubblicazione: (2025)
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
di: Pham, Long-Khanh, et al.
Pubblicazione: (2025)
di: Pham, Long-Khanh, et al.
Pubblicazione: (2025)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
di: Anh, Tran Nguyen, et al.
Pubblicazione: (2025)
di: Anh, Tran Nguyen, et al.
Pubblicazione: (2025)
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
di: Nguyen-Le, Hai-Son, et al.
Pubblicazione: (2026)
di: Nguyen-Le, Hai-Son, et al.
Pubblicazione: (2026)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
di: Jung, Jaemin, et al.
Pubblicazione: (2024)
di: Jung, Jaemin, et al.
Pubblicazione: (2024)
Acousto-optic reconstruction of exterior sound field based on concentric circle sampling with circular harmonic expansion
di: Nguyen, Phuc Duc, et al.
Pubblicazione: (2023)
di: Nguyen, Phuc Duc, et al.
Pubblicazione: (2023)
How to train your ears: Auditory-model emulation for large-dynamic-range inputs and mild-to-severe hearing losses
di: Leer, Peter, et al.
Pubblicazione: (2024)
di: Leer, Peter, et al.
Pubblicazione: (2024)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
di: Zhao, Shengkui, et al.
Pubblicazione: (2021)
di: Zhao, Shengkui, et al.
Pubblicazione: (2021)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2025)
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2025)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
BiosERC: Integrating Biography Speakers Supported by LLMs for ERC Tasks
di: Xue, Jieying, et al.
Pubblicazione: (2024)
di: Xue, Jieying, et al.
Pubblicazione: (2024)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2024)
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2024)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2024)
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
di: Nguyen, Tuan Nam, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan Nam, et al.
Pubblicazione: (2024)
Privacy Disclosure of Similarity Rank in Speech and Language Processing
di: Bäckström, Tom, et al.
Pubblicazione: (2025)
di: Bäckström, Tom, et al.
Pubblicazione: (2025)
AdaptVC: High Quality Voice Conversion with Adaptive Learning
di: Kim, Jaehun, et al.
Pubblicazione: (2025)
di: Kim, Jaehun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
di: Pham, The Hieu, et al.
Pubblicazione: (2025) -
PhoWhisper: Automatic Speech Recognition for Vietnamese
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024) -
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
di: Lam, Phat, et al.
Pubblicazione: (2024) -
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
di: Nguyen, Tuan, et al.
Pubblicazione: (2025) -
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
di: Nguyen, Tin, et al.
Pubblicazione: (2024)