Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
Fuente:
arXiv
Guardado en:
| Autores principales: | Cauzinille, Jules, Miron, Marius, Pietquin, Olivier, Hagiwara, Masato, Marxer, Ricard, Rey, Arnaud, Favre, Benoit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
por: Combes, Paolo, et al.
Publicado: (2025)
por: Combes, Paolo, et al.
Publicado: (2025)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026)
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
Enhancing Speaker Verification with Whispered Speech via Post-Processing
por: Gołębiowska, Magdalena, et al.
Publicado: (2026)
por: Gołębiowska, Magdalena, et al.
Publicado: (2026)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
por: Du, Wenzhang
Publicado: (2025)
por: Du, Wenzhang
Publicado: (2025)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
por: Chien, Sheng-You, et al.
Publicado: (2026)
por: Chien, Sheng-You, et al.
Publicado: (2026)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
por: Ismail, Saifelden M.
Publicado: (2025)
por: Ismail, Saifelden M.
Publicado: (2025)
How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios
por: Venkatesh, Satvik, et al.
Publicado: (2025)
por: Venkatesh, Satvik, et al.
Publicado: (2025)
Cepstral Smoothing of Binary Masks for Convolutive Blind Separation of Speech Mixtures
por: Missaoui, Ibrahim, et al.
Publicado: (2026)
por: Missaoui, Ibrahim, et al.
Publicado: (2026)
Quantum-Enhanced Analysis and Grading of Vocal Performance
por: Agarwal, Rohan
Publicado: (2025)
por: Agarwal, Rohan
Publicado: (2025)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
por: Khushiyant, et al.
Publicado: (2026)
por: Khushiyant, et al.
Publicado: (2026)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
por: Cao, Songjun, et al.
Publicado: (2026)
por: Cao, Songjun, et al.
Publicado: (2026)
RARR : Robust Real-World Activity Recognition with Vibration by Scavenging Near-Surface Audio Online
por: Lee, Dong Yoon, et al.
Publicado: (2025)
por: Lee, Dong Yoon, et al.
Publicado: (2025)
EMOVOME: A Dataset for Emotion Recognition in Spontaneous Real-Life Speech
por: Gómez-Zaragozá, Lucía, et al.
Publicado: (2024)
por: Gómez-Zaragozá, Lucía, et al.
Publicado: (2024)
Real-time Low-latency Music Source Separation using Hybrid Spectrogram-TasNet
por: Venkatesh, Satvik, et al.
Publicado: (2024)
por: Venkatesh, Satvik, et al.
Publicado: (2024)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
por: Aristorenas, Aris J.
Publicado: (2024)
por: Aristorenas, Aris J.
Publicado: (2024)
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
por: Cui, Hanfang, et al.
Publicado: (2025)
por: Cui, Hanfang, et al.
Publicado: (2025)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
por: Maity, Subhajit, et al.
Publicado: (2025)
por: Maity, Subhajit, et al.
Publicado: (2025)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
Parametric Digital Twins for Preserving Historic Buildings: A Case Study at Löfstad Castle in Östergötland, Sweden
por: Ni, Zhongjun, et al.
Publicado: (2024)
por: Ni, Zhongjun, et al.
Publicado: (2024)
Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models
por: Tralie, Christopher J., et al.
Publicado: (2024)
por: Tralie, Christopher J., et al.
Publicado: (2024)
RG-TTA: Regime-Guided Meta-Control for Test-Time Adaptation in Streaming Time Series
por: Kumar, Indar, et al.
Publicado: (2026)
por: Kumar, Indar, et al.
Publicado: (2026)
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
por: Husain, Jaavid Aktar, et al.
Publicado: (2026)
por: Husain, Jaavid Aktar, et al.
Publicado: (2026)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
por: Alamr, Meshal, et al.
Publicado: (2026)
por: Alamr, Meshal, et al.
Publicado: (2026)
Alternative Local Discriminant Bases Using Empirical Expectation and Variance Estimation
por: Fossgaard, Eirik
Publicado: (1999)
por: Fossgaard, Eirik
Publicado: (1999)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
por: Mukherjee, Debdeep, et al.
Publicado: (2025)
por: Mukherjee, Debdeep, et al.
Publicado: (2025)
Uncovering Population PK Covariates from VAE-Generated Latent Spaces
por: Perazzolo, Diego, et al.
Publicado: (2025)
por: Perazzolo, Diego, et al.
Publicado: (2025)
STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts
por: Opria, Joshua
Publicado: (2026)
por: Opria, Joshua
Publicado: (2026)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
por: Semenov, Andrei, et al.
Publicado: (2024)
por: Semenov, Andrei, et al.
Publicado: (2024)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
por: Bouzid, Hamza, et al.
Publicado: (2023)
por: Bouzid, Hamza, et al.
Publicado: (2023)
Apictorial Jigsaw Puzzle Reconstruction Based on Curve Matching via a Corotational Beam Spline
por: Orynyak, Igor, et al.
Publicado: (2025)
por: Orynyak, Igor, et al.
Publicado: (2025)
MEG-to-MEG Transfer Learning and Cross-Task Speech/Silence Detection with Limited Data
por: de Zuazo, Xabier, et al.
Publicado: (2026)
por: de Zuazo, Xabier, et al.
Publicado: (2026)
A Novel Schur-Decomposition-Based Weight Projection Method for Stable State-Space Neural-Network Architectures
por: Vanegas, Sergio, et al.
Publicado: (2026)
por: Vanegas, Sergio, et al.
Publicado: (2026)
Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
por: Sun, Ling, et al.
Publicado: (2025)
por: Sun, Ling, et al.
Publicado: (2025)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
por: Kozak, Nazar
Publicado: (2026)
por: Kozak, Nazar
Publicado: (2026)
VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
por: Basha, Maris, et al.
Publicado: (2025)
por: Basha, Maris, et al.
Publicado: (2025)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
por: Di Santi, Eduardo, et al.
Publicado: (2025)
por: Di Santi, Eduardo, et al.
Publicado: (2025)
Accurate typhoon intensity forecasts using a non-iterative spatiotemporal transformer model
por: Qu, Hongyu, et al.
Publicado: (2025)
por: Qu, Hongyu, et al.
Publicado: (2025)
Ejemplares similares
-
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
por: Combes, Paolo, et al.
Publicado: (2025) -
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026) -
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
por: Kim, Minu, et al.
Publicado: (2025) -
Enhancing Speaker Verification with Whispered Speech via Post-Processing
por: Gołębiowska, Magdalena, et al.
Publicado: (2026) -
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
por: Du, Wenzhang
Publicado: (2025)