Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yadav, Amit Kumar Singh, Xiang, Ziyue, Bhagtani, Kratika, Bestagini, Paolo, Tubaro, Stefano, Delp, Edward J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Comparative Analysis of ASR Methods for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
von: Bhagtani, Kratika, et al.
Veröffentlicht: (2024)
von: Bhagtani, Kratika, et al.
Veröffentlicht: (2024)
FairSSD: Understanding Bias in Synthetic Speech Detectors
von: Yadav, Amit Kumar Singh, et al.
Veröffentlicht: (2024)
von: Yadav, Amit Kumar Singh, et al.
Veröffentlicht: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Source Verification for Speech Deepfakes
von: Negroni, Viola, et al.
Veröffentlicht: (2025)
von: Negroni, Viola, et al.
Veröffentlicht: (2025)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2025)
von: Salvi, Davide, et al.
Veröffentlicht: (2025)
Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Analyzing the Impact of Splicing Artifacts in Partially Fake Speech Signals
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
Leveraging Mixture of Experts for Improved Speech Deepfake Detection
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
Audio Compression using Periodic Gabor with Biorthogonal Exchange: Implementation Using the Zak Transform
von: Alimi, Roger, et al.
Veröffentlicht: (2025)
von: Alimi, Roger, et al.
Veröffentlicht: (2025)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
von: Abbott, Leigh, et al.
Veröffentlicht: (2024)
von: Abbott, Leigh, et al.
Veröffentlicht: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
DiffAU: Diffusion-Based Ambisonics Upscaling
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
A Study on Speech Assessment with Visual Cues
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
Speech Watermarking with Discrete Intermediate Representations
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Speech dereverberation constrained on room impulse response characteristics
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Comparative Analysis of ASR Methods for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024) -
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
von: Bhagtani, Kratika, et al.
Veröffentlicht: (2024) -
FairSSD: Understanding Bias in Synthetic Speech Detectors
von: Yadav, Amit Kumar Singh, et al.
Veröffentlicht: (2024) -
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024) -
Source Verification for Speech Deepfakes
von: Negroni, Viola, et al.
Veröffentlicht: (2025)