Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Sarkar, Eklavya, -Doss, Mathew Magimai. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
by: Sarkar, Eklavya, et al.
Published: (2024)
by: Sarkar, Eklavya, et al.
Published: (2024)
Feature Representations for Automatic Meerkat Vocalization Classification
by: Mahmoud, Imen Ben, et al.
Published: (2024)
by: Mahmoud, Imen Ben, et al.
Published: (2024)
On feature representations for marmoset vocal communication analysis
by: Sarkar, Eklavya, et al.
Published: (2025)
by: Sarkar, Eklavya, et al.
Published: (2025)
Towards Leveraging Sequential Structure in Animal Vocalizations
by: Sarkar, Eklavya, et al.
Published: (2025)
by: Sarkar, Eklavya, et al.
Published: (2025)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
by: Hajal, Karl El, et al.
Published: (2025)
by: Hajal, Karl El, et al.
Published: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
by: Hajal, Karl El, et al.
Published: (2025)
by: Hajal, Karl El, et al.
Published: (2025)
kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
by: Hajal, Karl El, et al.
Published: (2024)
by: Hajal, Karl El, et al.
Published: (2024)
Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification
by: Sarkar, Eklavya, et al.
Published: (2026)
by: Sarkar, Eklavya, et al.
Published: (2026)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
by: Elbanna, Gasser, et al.
Published: (2024)
by: Elbanna, Gasser, et al.
Published: (2024)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
by: Zhang, Alice, et al.
Published: (2025)
by: Zhang, Alice, et al.
Published: (2025)
Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection
by: Moummad, Ilyass, et al.
Published: (2023)
by: Moummad, Ilyass, et al.
Published: (2023)
Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
by: Khorrami, Khazar, et al.
Published: (2023)
by: Khorrami, Khazar, et al.
Published: (2023)
Foundation Models for Bioacoustics -- a Comparative Review
by: Schwinger, Raphael, et al.
Published: (2025)
by: Schwinger, Raphael, et al.
Published: (2025)
Toward using Speech to Sense Student Emotion in Remote Learning Environments
by: Vyas, Sargam, et al.
Published: (2026)
by: Vyas, Sargam, et al.
Published: (2026)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
by: Fu, Yonggan, et al.
Published: (2022)
by: Fu, Yonggan, et al.
Published: (2022)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
by: Vaessen, Nik, et al.
Published: (2024)
by: Vaessen, Nik, et al.
Published: (2024)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
by: Caubrière, Antoine, et al.
Published: (2024)
by: Caubrière, Antoine, et al.
Published: (2024)
Perch 2.0: The Bittern Lesson for Bioacoustics
by: van Merriënboer, Bart, et al.
Published: (2025)
by: van Merriënboer, Bart, et al.
Published: (2025)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
by: Brima, Yusuf, et al.
Published: (2023)
by: Brima, Yusuf, et al.
Published: (2023)
Aggregation Strategies for Efficient Annotation of Bioacoustic Sound Events Using Active Learning
by: Lindholm, Richard, et al.
Published: (2025)
by: Lindholm, Richard, et al.
Published: (2025)
Towards interfacing large language models with ASR systems using confidence measures and prompting
by: Naderi, Maryam, et al.
Published: (2024)
by: Naderi, Maryam, et al.
Published: (2024)
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
by: Arcos-Holzinger, Sandra, et al.
Published: (2026)
by: Arcos-Holzinger, Sandra, et al.
Published: (2026)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
by: Menon, Aditya Srinivas, et al.
Published: (2026)
by: Menon, Aditya Srinivas, et al.
Published: (2026)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
by: Liu, Andy T., et al.
Published: (2024)
by: Liu, Andy T., et al.
Published: (2024)
SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
by: Shams, Siavash, et al.
Published: (2024)
by: Shams, Siavash, et al.
Published: (2024)
Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
by: Yang, Yudong, et al.
Published: (2024)
by: Yang, Yudong, et al.
Published: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
by: Lu, Yen-Ju, et al.
Published: (2024)
by: Lu, Yen-Ju, et al.
Published: (2024)
What Do Self-Supervised Speech Models Know About Words?
by: Pasad, Ankita, et al.
Published: (2023)
by: Pasad, Ankita, et al.
Published: (2023)
Weakly Supervised Detection and Temporal Localization of Whale Calls in Long-Duration Bioacoustic Data
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech
by: Samanta, Himadri S
Published: (2026)
by: Samanta, Himadri S
Published: (2026)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
by: Lovelace, Justin, et al.
Published: (2025)
by: Lovelace, Justin, et al.
Published: (2025)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
by: Sinha, Anshuman, et al.
Published: (2024)
by: Sinha, Anshuman, et al.
Published: (2024)
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
by: Kulkarni, Ajinkya, et al.
Published: (2025)
by: Kulkarni, Ajinkya, et al.
Published: (2025)
Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
by: Miara, Victor, et al.
Published: (2024)
by: Miara, Victor, et al.
Published: (2024)
Rethinking Mamba in Speech Processing by Self-Supervised Models
by: Zhang, Xiangyu, et al.
Published: (2024)
by: Zhang, Xiangyu, et al.
Published: (2024)
Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
by: Poncelet, Jakob, et al.
Published: (2021)
by: Poncelet, Jakob, et al.
Published: (2021)
An Empirical Analysis of Speech Self-Supervised Learning at Multiple Resolutions
by: Clark, Theo, et al.
Published: (2024)
by: Clark, Theo, et al.
Published: (2024)
Self-Supervised Speech Representations are More Phonetic than Semantic
by: Choi, Kwanghee, et al.
Published: (2024)
by: Choi, Kwanghee, et al.
Published: (2024)
Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
by: Razig, Amine, et al.
Published: (2025)
by: Razig, Amine, et al.
Published: (2025)
Similar Items
-
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
by: Sarkar, Eklavya, et al.
Published: (2024) -
Feature Representations for Automatic Meerkat Vocalization Classification
by: Mahmoud, Imen Ben, et al.
Published: (2024) -
On feature representations for marmoset vocal communication analysis
by: Sarkar, Eklavya, et al.
Published: (2025) -
Towards Leveraging Sequential Structure in Animal Vocalizations
by: Sarkar, Eklavya, et al.
Published: (2025) -
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
by: Hajal, Karl El, et al.
Published: (2025)