Forensic deepfake audio detection using segmental speech features
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Tianle, Sun, Chengzhe, Lyu, Siwei, Rose, Phil |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
par: Li, Xiang, et autres
Publié: (2024)
par: Li, Xiang, et autres
Publié: (2024)
A robust audio deepfake detection system via multi-view feature
par: Yang, Yujie, et autres
Publié: (2024)
par: Yang, Yujie, et autres
Publié: (2024)
Revisiting speech segmentation and lexicon learning with better features
par: Kamper, Herman, et autres
Publié: (2024)
par: Kamper, Herman, et autres
Publié: (2024)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
par: C, Anandh, et autres
Publié: (2025)
par: C, Anandh, et autres
Publié: (2025)
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
par: Pascu, Octavian, et autres
Publié: (2024)
par: Pascu, Octavian, et autres
Publié: (2024)
Echoes: A semantically-aligned music deepfake detection dataset
par: Pascu, Octavian, et autres
Publié: (2026)
par: Pascu, Octavian, et autres
Publié: (2026)
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
par: Yang, Chih-Kai, et autres
Publié: (2026)
par: Yang, Chih-Kai, et autres
Publié: (2026)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
par: Okocha, Chibuzor, et autres
Publié: (2025)
par: Okocha, Chibuzor, et autres
Publié: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
par: Yang, Tianle, et autres
Publié: (2026)
par: Yang, Tianle, et autres
Publié: (2026)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
par: Wang, Xiong, et autres
Publié: (2024)
par: Wang, Xiong, et autres
Publié: (2024)
NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
par: Mendes-Laureano, Janaína, et autres
Publié: (2024)
par: Mendes-Laureano, Janaína, et autres
Publié: (2024)
InstructAudio: Unified speech and music generation with natural language instruction
par: Qiang, Chunyu, et autres
Publié: (2025)
par: Qiang, Chunyu, et autres
Publié: (2025)
A unified front-end framework for English text-to-speech synthesis
par: Ying, Zelin, et autres
Publié: (2023)
par: Ying, Zelin, et autres
Publié: (2023)
Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks
par: Beguš, Gašper, et autres
Publié: (2023)
par: Beguš, Gašper, et autres
Publié: (2023)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
par: Wan, Xucheng, et autres
Publié: (2024)
par: Wan, Xucheng, et autres
Publié: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
par: Lonergan, Liam, et autres
Publié: (2024)
par: Lonergan, Liam, et autres
Publié: (2024)
learning discriminative features from spectrograms using center loss for speech emotion recognition
par: Dai, Dongyang, et autres
Publié: (2025)
par: Dai, Dongyang, et autres
Publié: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
par: Kloots, Marianne de Heer, et autres
Publié: (2025)
par: Kloots, Marianne de Heer, et autres
Publié: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
par: Gong, Rong, et autres
Publié: (2024)
par: Gong, Rong, et autres
Publié: (2024)
Controlling Surprisal in Music Generation via Information Content Curve Matching
par: Bjare, Mathias Rose, et autres
Publié: (2024)
par: Bjare, Mathias Rose, et autres
Publié: (2024)
Joint sentiment analysis of lyrics and audio in music
par: Schaab, Lea, et autres
Publié: (2024)
par: Schaab, Lea, et autres
Publié: (2024)
ADIFF: Explaining audio difference using natural language
par: Deshmukh, Soham, et autres
Publié: (2025)
par: Deshmukh, Soham, et autres
Publié: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
par: Gowda, Harshavardhana T., et autres
Publié: (2025)
par: Gowda, Harshavardhana T., et autres
Publié: (2025)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
par: Cui, Ziyun, et autres
Publié: (2023)
par: Cui, Ziyun, et autres
Publié: (2023)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
par: Wang, Shenran, et autres
Publié: (2025)
par: Wang, Shenran, et autres
Publié: (2025)
Generalizable speech deepfake detection via meta-learned LoRA
par: Laakkonen, Janne, et autres
Publié: (2025)
par: Laakkonen, Janne, et autres
Publié: (2025)
Discrete Audio Tokens: More Than a Survey!
par: Mousavi, Pooneh, et autres
Publié: (2025)
par: Mousavi, Pooneh, et autres
Publié: (2025)
Sustaining model performance for covid-19 detection from dynamic audio data: Development and evaluation of a comprehensive drift-adaptive framework
par: Ganitidis, Theofanis, et autres
Publié: (2024)
par: Ganitidis, Theofanis, et autres
Publié: (2024)
Hello-Chat: Towards Realistic Social Audio Interactions
par: Hou, Yueran, et autres
Publié: (2026)
par: Hou, Yueran, et autres
Publié: (2026)
Exploring Speech Pattern Disorders in Autism using Machine Learning
par: Hu, Chuanbo, et autres
Publié: (2024)
par: Hu, Chuanbo, et autres
Publié: (2024)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
par: Zhao, Qiuming, et autres
Publié: (2025)
par: Zhao, Qiuming, et autres
Publié: (2025)
Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
par: Wang, Zhenyu, et autres
Publié: (2024)
par: Wang, Zhenyu, et autres
Publié: (2024)
Roadmap towards Superhuman Speech Understanding using Large Language Models
par: Bu, Fan, et autres
Publié: (2024)
par: Bu, Fan, et autres
Publié: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
par: Araiza-Illan, Gloria, et autres
Publié: (2023)
par: Araiza-Illan, Gloria, et autres
Publié: (2023)
Moshi: a speech-text foundation model for real-time dialogue
par: Défossez, Alexandre, et autres
Publié: (2024)
par: Défossez, Alexandre, et autres
Publié: (2024)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
par: Deng, Yayue, et autres
Publié: (2025)
par: Deng, Yayue, et autres
Publié: (2025)
Mellow: a small audio language model for reasoning
par: Deshmukh, Soham, et autres
Publié: (2025)
par: Deshmukh, Soham, et autres
Publié: (2025)
GLAP: General contrastive audio-text pretraining across domains and languages
par: Dinkel, Heinrich, et autres
Publié: (2025)
par: Dinkel, Heinrich, et autres
Publié: (2025)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
par: Chen, Maximillian, et autres
Publié: (2024)
par: Chen, Maximillian, et autres
Publié: (2024)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
par: Yu, Wenyi, et autres
Publié: (2024)
par: Yu, Wenyi, et autres
Publié: (2024)
Documents similaires
-
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
par: Li, Xiang, et autres
Publié: (2024) -
A robust audio deepfake detection system via multi-view feature
par: Yang, Yujie, et autres
Publié: (2024) -
Revisiting speech segmentation and lexicon learning with better features
par: Kamper, Herman, et autres
Publié: (2024) -
Improving endpoint detection in end-to-end streaming ASR for conversational speech
par: C, Anandh, et autres
Publié: (2025) -
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
par: Pascu, Octavian, et autres
Publié: (2024)