Deep functional multiple index models with an application to SER
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saumard, Matthieu, Haj, Abir El, Napoleon, Thibault |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative AI-based data augmentation for improved bioacoustic classification in noisy environments
von: Gibbons, Anthony, et al.
Veröffentlicht: (2024)
von: Gibbons, Anthony, et al.
Veröffentlicht: (2024)
Pièces de viole des Cinq Livres and their statistical signatures: the musical work of Marin Marais and Jordi Savall
von: Lugo, Igor, et al.
Veröffentlicht: (2024)
von: Lugo, Igor, et al.
Veröffentlicht: (2024)
Complexity of frequency fluctuations and the interpretive style in the bass viola da gamba
von: Lugo, Igor, et al.
Veröffentlicht: (2025)
von: Lugo, Igor, et al.
Veröffentlicht: (2025)
The Rest is Silence: Leveraging Unseen Species Models for Computational Musicology
von: Moss, Fabian C., et al.
Veröffentlicht: (2025)
von: Moss, Fabian C., et al.
Veröffentlicht: (2025)
Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
von: Bonafos, Guillem, et al.
Veröffentlicht: (2024)
von: Bonafos, Guillem, et al.
Veröffentlicht: (2024)
Musical composition and 2D cellular automata based on music intervals
von: Lugo, Igor, et al.
Veröffentlicht: (2024)
von: Lugo, Igor, et al.
Veröffentlicht: (2024)
Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
von: Razig, Amine, et al.
Veröffentlicht: (2025)
von: Razig, Amine, et al.
Veröffentlicht: (2025)
Bayesian Restoration of Audio Degraded by Low-Frequency Pulses Modeled via Gaussian Process
von: de Carvalho, Hugo Tremonte, et al.
Veröffentlicht: (2020)
von: de Carvalho, Hugo Tremonte, et al.
Veröffentlicht: (2020)
THAI Speech Emotion Recognition (THAI-SER) corpus
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
Regularized autoregressive modeling and its application to audio signal reconstruction
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
Are audio DeepFake detection models polyglots?
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
von: Schäfer-Zimmermann, Julian C., et al.
Veröffentlicht: (2024)
von: Schäfer-Zimmermann, Julian C., et al.
Veröffentlicht: (2024)
Explainable anomaly detection for sound spectrograms using pooling statistics with quantile differences
von: Thewes, Nicolas, et al.
Veröffentlicht: (2025)
von: Thewes, Nicolas, et al.
Veröffentlicht: (2025)
Design framework for spherical microphone and loudspeaker arrays in a multiple-input multiple-output system
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Theory and investigation of acoustic multiple-input multiple-output systems based on spherical arrays in a room
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
Deep, data-driven modeling of room acoustics: literature review and research perspectives
von: van Waterschoot, Toon
Veröffentlicht: (2025)
von: van Waterschoot, Toon
Veröffentlicht: (2025)
DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
von: Kheir, Yassine El, et al.
Veröffentlicht: (2026)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2026)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
von: Pulikodan, Sujith, et al.
Veröffentlicht: (2025)
von: Pulikodan, Sujith, et al.
Veröffentlicht: (2025)
The importance of spatial and spectral information in multiple speaker tracking
von: Beit-On, Hanan, et al.
Veröffentlicht: (2024)
von: Beit-On, Hanan, et al.
Veröffentlicht: (2024)
AntiDeepFake: AI for Deep Fake Speech Recognition
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2024)
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2024)
On the application of Visibility Graphs in the Spectral Domain for Speaker Recognition
von: Bocaccio, Hernan, et al.
Veröffentlicht: (2025)
von: Bocaccio, Hernan, et al.
Veröffentlicht: (2025)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
Deep Learning for Personalized Binaural Audio Reproduction
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
Speech Representation Analysis based on Inter- and Intra-Model Similarities
von: Kheir, Yassine El, et al.
Veröffentlicht: (2024)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2024)
The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement
von: Leglaive, Simon, et al.
Veröffentlicht: (2023)
von: Leglaive, Simon, et al.
Veröffentlicht: (2023)
A Hierarchical Deep Learning Approach for Minority Instrument Detection
von: Sechet, Dylan, et al.
Veröffentlicht: (2025)
von: Sechet, Dylan, et al.
Veröffentlicht: (2025)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
Deep Speech Synthesis from Multimodal Articulatory Representations
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
Onset and offset weighted loss function for sound event detection
von: Song, Tao
Veröffentlicht: (2024)
von: Song, Tao
Veröffentlicht: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)
Frequency Tracking Features for Data-Efficient Deep Siren Identification
von: Damiano, Stefano, et al.
Veröffentlicht: (2024)
von: Damiano, Stefano, et al.
Veröffentlicht: (2024)
Estimating the Number and Locations of Boundaries in Reverberant Environments with Deep Learning
von: Arikan, Toros, et al.
Veröffentlicht: (2024)
von: Arikan, Toros, et al.
Veröffentlicht: (2024)
Room Impulse Responses help attackers to evade Deep Fake Detection
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
Deep learning based spatial aliasing reduction in beamforming for audio capture
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generative AI-based data augmentation for improved bioacoustic classification in noisy environments
von: Gibbons, Anthony, et al.
Veröffentlicht: (2024) -
Pièces de viole des Cinq Livres and their statistical signatures: the musical work of Marin Marais and Jordi Savall
von: Lugo, Igor, et al.
Veröffentlicht: (2024) -
Complexity of frequency fluctuations and the interpretive style in the bass viola da gamba
von: Lugo, Igor, et al.
Veröffentlicht: (2025) -
The Rest is Silence: Leveraging Unseen Species Models for Computational Musicology
von: Moss, Fabian C., et al.
Veröffentlicht: (2025) -
Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
von: Bonafos, Guillem, et al.
Veröffentlicht: (2024)