Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Silaev, Mikhail, Drossos, Konstantinos, Virtanen, Tuomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
von: Neri, Michael, et al.
Veröffentlicht: (2026)
von: Neri, Michael, et al.
Veröffentlicht: (2026)
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026)
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026)
Moving Speaker Separation via Parallel Spectral-Spatial Processing
von: Wang, Yuzhu, et al.
Veröffentlicht: (2026)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2026)
Multi-Utterance Speech Separation and Association Trained on Short Segments
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
von: Drossos, Konstantinos, et al.
Veröffentlicht: (2025)
von: Drossos, Konstantinos, et al.
Veröffentlicht: (2025)
Acoustic Simulation Framework for Multi-channel Replay Speech Detection
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026)
von: Luong, Diep, et al.
Veröffentlicht: (2026)
Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
von: Luong, Diep, et al.
Veröffentlicht: (2025)
von: Luong, Diep, et al.
Veröffentlicht: (2025)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
von: Steckel, Jan, et al.
Veröffentlicht: (2024)
von: Steckel, Jan, et al.
Veröffentlicht: (2024)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
Representation Learning for Audio Privacy Preservation using Source Separation and Robust Adversarial Learning
von: Luong, Diep, et al.
Veröffentlicht: (2023)
von: Luong, Diep, et al.
Veröffentlicht: (2023)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications
von: Diaz-Guerra, David, et al.
Veröffentlicht: (2023)
von: Diaz-Guerra, David, et al.
Veröffentlicht: (2023)
Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
von: Voran, Stephen D.
Veröffentlicht: (2024)
von: Voran, Stephen D.
Veröffentlicht: (2024)
AI-Generated Music Detection in Broadcast Monitoring
von: López-Ayala, David, et al.
Veröffentlicht: (2026)
von: López-Ayala, David, et al.
Veröffentlicht: (2026)
Speech Enhancement Based on Drifting Models
von: Xu, Liang, et al.
Veröffentlicht: (2026)
von: Xu, Liang, et al.
Veröffentlicht: (2026)
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2026)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2026)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
Wavetable Synthesis Using CVAE for Timbre Control Based on Semantic Label
von: Yutani, Tsugumasa, et al.
Veröffentlicht: (2024)
von: Yutani, Tsugumasa, et al.
Veröffentlicht: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
von: Latifi, Seyed Amir, et al.
Veröffentlicht: (2024)
von: Latifi, Seyed Amir, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
von: Neri, Michael, et al.
Veröffentlicht: (2026) -
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026) -
Moving Speaker Separation via Parallel Spectral-Spatial Processing
von: Wang, Yuzhu, et al.
Veröffentlicht: (2026) -
Multi-Utterance Speech Separation and Association Trained on Short Segments
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025) -
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)