Gespeichert in:
| Hauptverfasser: | Yang, Tianle, Sun, Chengzhe, Lyu, Siwei, Rose, Phil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.13847 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
A robust audio deepfake detection system via multi-view feature
von: Yang, Yujie, et al.
Veröffentlicht: (2024)
von: Yang, Yujie, et al.
Veröffentlicht: (2024)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
von: C, Anandh, et al.
Veröffentlicht: (2025)
von: C, Anandh, et al.
Veröffentlicht: (2025)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
von: Pascu, Octavian, et al.
Veröffentlicht: (2024)
von: Pascu, Octavian, et al.
Veröffentlicht: (2024)
Echoes: A semantically-aligned music deepfake detection dataset
von: Pascu, Octavian, et al.
Veröffentlicht: (2026)
von: Pascu, Octavian, et al.
Veröffentlicht: (2026)
Acoustic and perceptual differences between standard and accented speech and their voice clones
von: Yang, Tianle, et al.
Veröffentlicht: (2026)
von: Yang, Tianle, et al.
Veröffentlicht: (2026)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
von: Mendes-Laureano, Janaína, et al.
Veröffentlicht: (2024)
von: Mendes-Laureano, Janaína, et al.
Veröffentlicht: (2024)
InstructAudio: Unified speech and music generation with natural language instruction
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
A unified front-end framework for English text-to-speech synthesis
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks
von: Beguš, Gašper, et al.
Veröffentlicht: (2023)
von: Beguš, Gašper, et al.
Veröffentlicht: (2023)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
Joint sentiment analysis of lyrics and audio in music
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
learning discriminative features from spectrograms using center loss for speech emotion recognition
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
Controlling Surprisal in Music Generation via Information Content Curve Matching
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2024)
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
ADIFF: Explaining audio difference using natural language
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
Moshi: a speech-text foundation model for real-time dialogue
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
Discrete Audio Tokens: More Than a Survey!
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
von: Cui, Ziyun, et al.
Veröffentlicht: (2023)
von: Cui, Ziyun, et al.
Veröffentlicht: (2023)
Hello-Chat: Towards Realistic Social Audio Interactions
von: Hou, Yueran, et al.
Veröffentlicht: (2026)
von: Hou, Yueran, et al.
Veröffentlicht: (2026)
Generalizable speech deepfake detection via meta-learned LoRA
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
Exploring Speech Pattern Disorders in Autism using Machine Learning
von: Hu, Chuanbo, et al.
Veröffentlicht: (2024)
von: Hu, Chuanbo, et al.
Veröffentlicht: (2024)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
Roadmap towards Superhuman Speech Understanding using Large Language Models
von: Bu, Fan, et al.
Veröffentlicht: (2024)
von: Bu, Fan, et al.
Veröffentlicht: (2024)
LLM-Driven Multimodal Opinion Expression Identification
von: Jia, Bonian, et al.
Veröffentlicht: (2024)
von: Jia, Bonian, et al.
Veröffentlicht: (2024)
Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
von: Manakul, Potsawee, et al.
Veröffentlicht: (2024)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2024)
Thunder : Unified Regression-Diffusion Speech Enhancement with a Single Reverse Step using Brownian Bridge
von: Trachu, Thanapat, et al.
Veröffentlicht: (2024)
von: Trachu, Thanapat, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
von: Li, Xiang, et al.
Veröffentlicht: (2024) -
A robust audio deepfake detection system via multi-view feature
von: Yang, Yujie, et al.
Veröffentlicht: (2024) -
Improving endpoint detection in end-to-end streaming ASR for conversational speech
von: C, Anandh, et al.
Veröffentlicht: (2025) -
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024) -
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)