Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
Fuente:
arXiv
Guardado en:
| Autores principales: | Bonafos, Guillem, Bourot, Clara, Pudlo, Pierre, Freyermuth, Jean-Marc, Reboul, Laurence, Tronçon, Samuel, Rey, Arnaud |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On feature representations for marmoset vocal communication analysis
por: Sarkar, Eklavya, et al.
Publicado: (2025)
por: Sarkar, Eklavya, et al.
Publicado: (2025)
Interfacing with history: Curating with audio augmented objects
por: Cliffe, Laurence
Publicado: (2024)
por: Cliffe, Laurence
Publicado: (2024)
Benchmarking multi-component signal processing methods in the time-frequency plane
por: Miramont, Juan M., et al.
Publicado: (2024)
por: Miramont, Juan M., et al.
Publicado: (2024)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
por: Zhou, Wangjin, et al.
Publicado: (2024)
por: Zhou, Wangjin, et al.
Publicado: (2024)
Improving acoustic drone detection generalization through pretraining and data augmentation
por: Reuter, Paul M., et al.
Publicado: (2026)
por: Reuter, Paul M., et al.
Publicado: (2026)
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
por: Lavechin, Marvin, et al.
Publicado: (2025)
por: Lavechin, Marvin, et al.
Publicado: (2025)
Feature Selection via Graph Topology Inference for Soundscape Emotion Recognition
por: Rey, Samuel, et al.
Publicado: (2025)
por: Rey, Samuel, et al.
Publicado: (2025)
Developing vocal system impaired patient-aimed voice quality assessment approach using ASR representation-included multiple features
por: Dang, Shaoxiang, et al.
Publicado: (2024)
por: Dang, Shaoxiang, et al.
Publicado: (2024)
Sample adaptive data augmentation with progressive scheduling
por: Lu, Hongxuan, et al.
Publicado: (2024)
por: Lu, Hongxuan, et al.
Publicado: (2024)
Blind estimation of audio effects using an auto-encoder approach and differentiable digital signal processing
por: Peladeau, Côme, et al.
Publicado: (2023)
por: Peladeau, Côme, et al.
Publicado: (2023)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
por: Li, Junjie, et al.
Publicado: (2024)
por: Li, Junjie, et al.
Publicado: (2024)
Privacy-oriented manipulation of speaker representations
por: Teixeira, Francisco, et al.
Publicado: (2023)
por: Teixeira, Francisco, et al.
Publicado: (2023)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
por: Slabbert, Danel, et al.
Publicado: (2025)
por: Slabbert, Danel, et al.
Publicado: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
por: An, Keyu, et al.
Publicado: (2025)
por: An, Keyu, et al.
Publicado: (2025)
Trusted Fake Audio Detection Based on Dirichlet Distribution
por: Ding, Chi, et al.
Publicado: (2025)
por: Ding, Chi, et al.
Publicado: (2025)
Réduire le bruit grâce à la réalité augmentée sonore -- Auditory Concealer
por: Boukhemia, Clara
Publicado: (2025)
por: Boukhemia, Clara
Publicado: (2025)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
por: Ronchini, Francesca, et al.
Publicado: (2020)
por: Ronchini, Francesca, et al.
Publicado: (2020)
Utilizing synthetic training data for the supervised classification of rat ultrasonic vocalizations
por: Scott, K. Jack, et al.
Publicado: (2023)
por: Scott, K. Jack, et al.
Publicado: (2023)
Hierarchical speaker representation for target speaker extraction
por: He, Shulin, et al.
Publicado: (2022)
por: He, Shulin, et al.
Publicado: (2022)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
por: Ueda, Lucas H., et al.
Publicado: (2026)
por: Ueda, Lucas H., et al.
Publicado: (2026)
Tweaking autoregressive methods for inpainting of gaps in audio signals
por: Mokrý, Ondřej, et al.
Publicado: (2024)
por: Mokrý, Ondřej, et al.
Publicado: (2024)
Comparison of fundamental frequency estimators with subharmonic voice signals
por: Ikuma, Takeshi, et al.
Publicado: (2025)
por: Ikuma, Takeshi, et al.
Publicado: (2025)
Generative AI-based data augmentation for improved bioacoustic classification in noisy environments
por: Gibbons, Anthony, et al.
Publicado: (2024)
por: Gibbons, Anthony, et al.
Publicado: (2024)
Speech enhancement deep-learning architecture for efficient edge processing
por: Pal, Monisankha, et al.
Publicado: (2024)
por: Pal, Monisankha, et al.
Publicado: (2024)
Binaural rendering from microphone array signals of arbitrary geometry
por: Iijima, Naoto, et al.
Publicado: (2021)
por: Iijima, Naoto, et al.
Publicado: (2021)
Regularized autoregressive modeling and its application to audio signal reconstruction
por: Mokrý, Ondřej, et al.
Publicado: (2024)
por: Mokrý, Ondřej, et al.
Publicado: (2024)
LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
por: Han, Changheon, et al.
Publicado: (2025)
por: Han, Changheon, et al.
Publicado: (2025)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
por: Araz, R. Oguz, et al.
Publicado: (2025)
por: Araz, R. Oguz, et al.
Publicado: (2025)
Source Separation by Flow Matching
por: Scheibler, Robin, et al.
Publicado: (2025)
por: Scheibler, Robin, et al.
Publicado: (2025)
Facilitating deep acoustic phenotyping: A basic coding scheme of infant vocalisations preluding computational analysis, machine learning and clinical reasoning
por: Kulvicius, Tomas, et al.
Publicado: (2023)
por: Kulvicius, Tomas, et al.
Publicado: (2023)
AxLSTMs: learning self-supervised audio representations with xLSTMs
por: Yadav, Sarthak, et al.
Publicado: (2024)
por: Yadav, Sarthak, et al.
Publicado: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
por: Pascu, Octavian, et al.
Publicado: (2023)
por: Pascu, Octavian, et al.
Publicado: (2023)
Acoustic-to-articulatory Inversion of the Complete Vocal Tract from RT-MRI with Various Audio Embeddings and Dataset Sizes
por: Azzouz, Sofiane, et al.
Publicado: (2026)
por: Azzouz, Sofiane, et al.
Publicado: (2026)
Complete reconstruction of the tongue contour through acoustic to articulatory inversion using real-time MRI data
por: Azzouz, Sofiane, et al.
Publicado: (2024)
por: Azzouz, Sofiane, et al.
Publicado: (2024)
Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model
por: Azzouz, Sofiane, et al.
Publicado: (2026)
por: Azzouz, Sofiane, et al.
Publicado: (2026)
Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
por: Azzouz, Sofiane, et al.
Publicado: (2026)
por: Azzouz, Sofiane, et al.
Publicado: (2026)
Reconstruction of the Complete Vocal Tract Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data
por: Azzouz, Sofiane, et al.
Publicado: (2025)
por: Azzouz, Sofiane, et al.
Publicado: (2025)
Enhancing the analysis of murine neonatal ultrasonic vocalizations: Development, evaluation, and application of different mathematical models
por: Herdt, Rudolf, et al.
Publicado: (2024)
por: Herdt, Rudolf, et al.
Publicado: (2024)
The role of direct sound spherical harmonics representation in externalization using binaural reproduction
por: Miller, Eran, et al.
Publicado: (2024)
por: Miller, Eran, et al.
Publicado: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024)
por: Watanabe, Aya, et al.
Publicado: (2024)
Ejemplares similares
-
On feature representations for marmoset vocal communication analysis
por: Sarkar, Eklavya, et al.
Publicado: (2025) -
Interfacing with history: Curating with audio augmented objects
por: Cliffe, Laurence
Publicado: (2024) -
Benchmarking multi-component signal processing methods in the time-frequency plane
por: Miramont, Juan M., et al.
Publicado: (2024) -
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
por: Zhou, Wangjin, et al.
Publicado: (2024) -
Improving acoustic drone detection generalization through pretraining and data augmentation
por: Reuter, Paul M., et al.
Publicado: (2026)