Gespeichert in:
| Hauptverfasser: | Wagner, Dominik, Churchill, Alexander, Sigtia, Siddharth, Georgiou, Panayiotis, Mirsamadi, Matt, Mishra, Aarshee, Marchi, Erik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2403.14438 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
von: Palaskar, Shruti, et al.
Veröffentlicht: (2024)
von: Palaskar, Shruti, et al.
Veröffentlicht: (2024)
Towards Explainable Spoofed Speech Attribution and Detection:a Probabilistic Approach for Characterizing Speech Synthesizer Components
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2025)
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2025)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
Adaptive Knowledge Distillation for Device-Directed Speech Detection
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
von: Ognjen, et al.
Veröffentlicht: (2024)
von: Ognjen, et al.
Veröffentlicht: (2024)
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
von: Xie, Jiamin, et al.
Veröffentlicht: (2025)
von: Xie, Jiamin, et al.
Veröffentlicht: (2025)
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
von: N, Rishith Sadashiv T, et al.
Veröffentlicht: (2025)
von: N, Rishith Sadashiv T, et al.
Veröffentlicht: (2025)
Predicting Cognitive Decline: A Multimodal AI Approach to Dementia Screening from Speech
von: Chi, Lei, et al.
Veröffentlicht: (2025)
von: Chi, Lei, et al.
Veröffentlicht: (2025)
A Survey on Speech Large Language Models for Understanding
von: Peng, Jing, et al.
Veröffentlicht: (2024)
von: Peng, Jing, et al.
Veröffentlicht: (2024)
Spatial Audio Processing with Large Language Model on Wearable Devices
von: Mishra, Ayushi, et al.
Veröffentlicht: (2025)
von: Mishra, Ayushi, et al.
Veröffentlicht: (2025)
Leveraging Cascaded Binary Classification and Multimodal Fusion for Dementia Detection through Spontaneous Speech
von: Liu, Yin-Long, et al.
Veröffentlicht: (2025)
von: Liu, Yin-Long, et al.
Veröffentlicht: (2025)
Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
von: Chhibber, Manasi, et al.
Veröffentlicht: (2025)
von: Chhibber, Manasi, et al.
Veröffentlicht: (2025)
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
von: Chinchmalatpure, Prajwal, et al.
Veröffentlicht: (2025)
von: Chinchmalatpure, Prajwal, et al.
Veröffentlicht: (2025)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
Group Relative Policy Optimization for Text-to-Speech with Large Language Models
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
Towards Frame-level Quality Predictions of Synthetic Speech
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2025)
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2025)
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
Direct Preference Optimization for Speech Autoregressive Diffusion Models
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2026)
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2026)
Device Feature based on Graph Fourier Transformation with Logarithmic Processing For Detection of Replay Speech Attacks
von: He, Mingrui, et al.
Veröffentlicht: (2024)
von: He, Mingrui, et al.
Veröffentlicht: (2024)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
von: Ren, Bo, et al.
Veröffentlicht: (2026)
von: Ren, Bo, et al.
Veröffentlicht: (2026)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and Machine Learning Approaches
von: Pierotti, Francesco, et al.
Veröffentlicht: (2025)
von: Pierotti, Francesco, et al.
Veröffentlicht: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Spoofing-Robust Speaker Verification Using Parallel Embedding Fusion: BTU Speech Group's Approach for ASVspoof5 Challenge
von: Kurnaz, Oğuzhan, et al.
Veröffentlicht: (2024)
von: Kurnaz, Oğuzhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
von: Wagner, Dominik, et al.
Veröffentlicht: (2025) -
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024) -
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
von: Palaskar, Shruti, et al.
Veröffentlicht: (2024) -
Towards Explainable Spoofed Speech Attribution and Detection:a Probabilistic Approach for Characterizing Speech Synthesizer Components
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2025) -
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)