Speech transformer models for extracting information from baby cries
Fuente:
arXiv
Guardado en:
| Autores principales: | Bonafos, Guillem, Rouch, Jéremy, Lego, Lény, Reby, David, Patural, Hugues, Mathevon, Nicolas, Emonet, Rémy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
por: Bonafos, Guillem, et al.
Publicado: (2024)
por: Bonafos, Guillem, et al.
Publicado: (2024)
Acoustic evaluation of a neural network dedicated to the detection of animal vocalisations
por: Rouch, Jérémy, et al.
Publicado: (2025)
por: Rouch, Jérémy, et al.
Publicado: (2025)
Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
por: Razig, Amine, et al.
Publicado: (2025)
por: Razig, Amine, et al.
Publicado: (2025)
Bayesian Restoration of Audio Degraded by Low-Frequency Pulses Modeled via Gaussian Process
por: de Carvalho, Hugo Tremonte, et al.
Publicado: (2020)
por: de Carvalho, Hugo Tremonte, et al.
Publicado: (2020)
Ellipsoid fitting with the Cayley transform
por: Melikechi, Omar, et al.
Publicado: (2023)
por: Melikechi, Omar, et al.
Publicado: (2023)
Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to Bahasa Melayu
por: Lam, Bhan, et al.
Publicado: (2022)
por: Lam, Bhan, et al.
Publicado: (2022)
Come Together: Analyzing Popular Songs Through Statistical Embeddings
por: Mallory, Matthew Esmaili, et al.
Publicado: (2026)
por: Mallory, Matthew Esmaili, et al.
Publicado: (2026)
Automobile demand forecasting: Spatiotemporal and hierarchical modeling, life cycle dynamics, and user-generated online information
por: Nahrendorf, Tom, et al.
Publicado: (2025)
por: Nahrendorf, Tom, et al.
Publicado: (2025)
Masked Autoencoders as Universal Speech Enhancer
por: Rajagopalan, Rajalaxmi, et al.
Publicado: (2026)
por: Rajagopalan, Rajalaxmi, et al.
Publicado: (2026)
Deep functional multiple index models with an application to SER
por: Saumard, Matthieu, et al.
Publicado: (2024)
por: Saumard, Matthieu, et al.
Publicado: (2024)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
por: Jiang, Xilin, et al.
Publicado: (2024)
por: Jiang, Xilin, et al.
Publicado: (2024)
Koopman Regularized Deep Speech Disentanglement for Speaker Verification
por: Chazaridis, Nikos, et al.
Publicado: (2026)
por: Chazaridis, Nikos, et al.
Publicado: (2026)
Assessing the Impact of Speaker Identity in Speech Spoofing Detection
por: Dao, Anh-Tuan, et al.
Publicado: (2026)
por: Dao, Anh-Tuan, et al.
Publicado: (2026)
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
por: Zhang, Ziqian, et al.
Publicado: (2025)
por: Zhang, Ziqian, et al.
Publicado: (2025)
IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
por: Padhya, Dinanath, et al.
Publicado: (2026)
por: Padhya, Dinanath, et al.
Publicado: (2026)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
A Semi-Supervised Framework for Speech Confidence Detection using Whisper
por: Wynn, Adam, et al.
Publicado: (2026)
por: Wynn, Adam, et al.
Publicado: (2026)
Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments
por: Ramamoorthy, Arnav
Publicado: (2025)
por: Ramamoorthy, Arnav
Publicado: (2025)
Investigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments
por: Anacin, et al.
Publicado: (2026)
por: Anacin, et al.
Publicado: (2026)
Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model
por: Ahn, Chung-Soo, et al.
Publicado: (2025)
por: Ahn, Chung-Soo, et al.
Publicado: (2025)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
por: Feng, Pengchao, et al.
Publicado: (2025)
por: Feng, Pengchao, et al.
Publicado: (2025)
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
por: Hirschkind, Nameer, et al.
Publicado: (2024)
por: Hirschkind, Nameer, et al.
Publicado: (2024)
EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
por: Muppidi, Akshay, et al.
Publicado: (2025)
por: Muppidi, Akshay, et al.
Publicado: (2025)
PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
por: Pahar, Madhurananda, et al.
Publicado: (2026)
por: Pahar, Madhurananda, et al.
Publicado: (2026)
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
por: Adnan, Tariq, et al.
Publicado: (2024)
por: Adnan, Tariq, et al.
Publicado: (2024)
mmWave Radar Aware Dual-Conditioned GAN for Speech Reconstruction of Signals With Low SNR
por: Karani, Jash, et al.
Publicado: (2026)
por: Karani, Jash, et al.
Publicado: (2026)
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
por: Justus, Aju Ani, et al.
Publicado: (2026)
por: Justus, Aju Ani, et al.
Publicado: (2026)
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
por: Kumar, Sahil, et al.
Publicado: (2026)
por: Kumar, Sahil, et al.
Publicado: (2026)
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
por: Wang, Jiaqi, et al.
Publicado: (2025)
por: Wang, Jiaqi, et al.
Publicado: (2025)
SpectroFusion-ViT: A Lightweight Transformer for Speech Emotion Recognition Using Harmonic Mel-Chroma Fusion
por: Ahmed, Faria, et al.
Publicado: (2026)
por: Ahmed, Faria, et al.
Publicado: (2026)
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
por: Borisov, Maksim, et al.
Publicado: (2025)
por: Borisov, Maksim, et al.
Publicado: (2025)
The added value for MRI radiomics and deep-learning for glioblastoma prognostication compared to clinical and molecular information
por: Abler, D., et al.
Publicado: (2025)
por: Abler, D., et al.
Publicado: (2025)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
por: Cuervo, Santiago, et al.
Publicado: (2024)
por: Cuervo, Santiago, et al.
Publicado: (2024)
An unified approach to link prediction in collaboration networks
por: Sosa, Juan, et al.
Publicado: (2024)
por: Sosa, Juan, et al.
Publicado: (2024)
Variational Bayes Portfolio Construction
por: Nguyen, Nicolas, et al.
Publicado: (2024)
por: Nguyen, Nicolas, et al.
Publicado: (2024)
Speech to Speech Synthesis for Voice Impersonation
por: Johnson, Bjorn, et al.
Publicado: (2026)
por: Johnson, Bjorn, et al.
Publicado: (2026)
Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2
por: Mohanty, Suvendu Sekhar
Publicado: (2026)
por: Mohanty, Suvendu Sekhar
Publicado: (2026)
Foundation for unbiased cross-validation of spatio-temporal models for species distribution modeling
por: Koldasbayeva, Diana, et al.
Publicado: (2025)
por: Koldasbayeva, Diana, et al.
Publicado: (2025)
Split and Conquer Partial Deepfake Speech
por: Rimon, Inbal, et al.
Publicado: (2026)
por: Rimon, Inbal, et al.
Publicado: (2026)
Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds
por: Moummad, Ilyass, et al.
Publicado: (2024)
por: Moummad, Ilyass, et al.
Publicado: (2024)
Ejemplares similares
-
Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
por: Bonafos, Guillem, et al.
Publicado: (2024) -
Acoustic evaluation of a neural network dedicated to the detection of animal vocalisations
por: Rouch, Jérémy, et al.
Publicado: (2025) -
Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
por: Razig, Amine, et al.
Publicado: (2025) -
Bayesian Restoration of Audio Degraded by Low-Frequency Pulses Modeled via Gaussian Process
por: de Carvalho, Hugo Tremonte, et al.
Publicado: (2020) -
Ellipsoid fitting with the Cayley transform
por: Melikechi, Omar, et al.
Publicado: (2023)