Salvato in:
| Autori principali: | Phuong, Tuan Dat, Truong, Duc-Tuan, Hoang, Long-Vu, Thu, Trang Nguyen Thi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.04702 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2024)
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2024)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2025)
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2025)
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2025)
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2025)
Continuous Learning of Transformer-based Audio Deepfake Detection
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
di: Chu, The Chuong, et al.
Pubblicazione: (2025)
di: Chu, The Chuong, et al.
Pubblicazione: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
Room Impulse Responses help attackers to evade Deep Fake Detection
di: Luong, Hieu-Thi, et al.
Pubblicazione: (2024)
di: Luong, Hieu-Thi, et al.
Pubblicazione: (2024)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Environmental Sound Deepfake Detection Using Deep-Learning Framework
di: Pham, Lam, et al.
Pubblicazione: (2026)
di: Pham, Lam, et al.
Pubblicazione: (2026)
Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators
di: Pham, Lam, et al.
Pubblicazione: (2026)
di: Pham, Lam, et al.
Pubblicazione: (2026)
Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing
di: Liu, Tianchi, et al.
Pubblicazione: (2025)
di: Liu, Tianchi, et al.
Pubblicazione: (2025)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
di: Vu, Tung, et al.
Pubblicazione: (2026)
di: Vu, Tung, et al.
Pubblicazione: (2026)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
Attention-based Mixture of Experts for Robust Speech Deepfake Detection
di: Negroni, Viola, et al.
Pubblicazione: (2025)
di: Negroni, Viola, et al.
Pubblicazione: (2025)
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
di: Tu, Huu Tuong, et al.
Pubblicazione: (2025)
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
di: Ho, Tuan Vu, et al.
Pubblicazione: (2024)
di: Ho, Tuan Vu, et al.
Pubblicazione: (2024)
Assessing the Impact of Speaker Identity in Speech Spoofing Detection
di: Dao, Anh-Tuan, et al.
Pubblicazione: (2026)
di: Dao, Anh-Tuan, et al.
Pubblicazione: (2026)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
di: Dao, Alan, et al.
Pubblicazione: (2025)
di: Dao, Alan, et al.
Pubblicazione: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Frame-level Temporal Difference Learning for Partial Deepfake Speech Detection
di: Li, Menglu, et al.
Pubblicazione: (2025)
di: Li, Menglu, et al.
Pubblicazione: (2025)
Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling
di: Negroni, Viola, et al.
Pubblicazione: (2026)
di: Negroni, Viola, et al.
Pubblicazione: (2026)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
di: Li, Junjie, et al.
Pubblicazione: (2025)
di: Li, Junjie, et al.
Pubblicazione: (2025)
Real-time Speech Summarization for Medical Conversations
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
di: Xie, Yuankun, et al.
Pubblicazione: (2025)
di: Xie, Yuankun, et al.
Pubblicazione: (2025)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Towards Scalable AASIST: Refining Graph Attention for Speech Deepfake Detection
di: Viakhirev, Ivan, et al.
Pubblicazione: (2025)
di: Viakhirev, Ivan, et al.
Pubblicazione: (2025)
Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2023)
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2023)
SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
di: Zhu, Yi, et al.
Pubblicazione: (2024)
di: Zhu, Yi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025) -
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
di: Truong, Duc-Tuan, et al.
Pubblicazione: (2024) -
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
di: Vu, Hoang Long, et al.
Pubblicazione: (2024) -
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025) -
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)