Enabling automatic transcription of child-centered audio recordings from real-world environments
Fuente:
arXiv
Saved in:
| Main Authors: | Kocharov, Daniil, Räsänen, Okko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
by: Räsänen, Okko, et al.
Published: (2024)
by: Räsänen, Okko, et al.
Published: (2024)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
by: Padhya, Dinanath, et al.
Published: (2026)
by: Padhya, Dinanath, et al.
Published: (2026)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
by: Ferreira, Alexandre R., et al.
Published: (2023)
by: Ferreira, Alexandre R., et al.
Published: (2023)
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
by: Pahar, Madhurananda, et al.
Published: (2025)
by: Pahar, Madhurananda, et al.
Published: (2025)
Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
by: Zehren, Mickaël, et al.
Published: (2024)
by: Zehren, Mickaël, et al.
Published: (2024)
GRAM: Spatial general-purpose audio representations for real-world environments
by: Yuksel, Goksenin, et al.
Published: (2026)
by: Yuksel, Goksenin, et al.
Published: (2026)
Training chord recognition models on artificially generated audio
by: Majchrzak, Martyna, et al.
Published: (2025)
by: Majchrzak, Martyna, et al.
Published: (2025)
Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
by: Riley, Xavier, et al.
Published: (2024)
by: Riley, Xavier, et al.
Published: (2024)
Context-aware child-directed speech detection from long-form recordings
by: Charlot, Théo, et al.
Published: (2026)
by: Charlot, Théo, et al.
Published: (2026)
Better audio representations are more brain-like: linking model-brain alignment with performance in downstream auditory tasks
by: Pepino, Leonardo, et al.
Published: (2025)
by: Pepino, Leonardo, et al.
Published: (2025)
The silence of the weights: a structural pruning strategy for attention-based audio signal architectures with second order metrics
by: Diecidue, Andrea, et al.
Published: (2025)
by: Diecidue, Andrea, et al.
Published: (2025)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
Transformation of audio embeddings into interpretable, concept-based representations
by: Zhang, Alice, et al.
Published: (2025)
by: Zhang, Alice, et al.
Published: (2025)
Towards generalizing deep-audio fake detection networks
by: Gasenzer, Konstantin, et al.
Published: (2023)
by: Gasenzer, Konstantin, et al.
Published: (2023)
Testing chatbots on the creation of encoders for audio conditioned image generation
by: León, Jorge E., et al.
Published: (2025)
by: León, Jorge E., et al.
Published: (2025)
Unsupervised outlier detection to improve bird audio dataset labels
by: Collins, Bruce
Published: (2025)
by: Collins, Bruce
Published: (2025)
Combining audio control and style transfer using latent diffusion
by: Demerlé, Nils, et al.
Published: (2024)
by: Demerlé, Nils, et al.
Published: (2024)
Versatile audio-visual learning for emotion recognition
by: Goncalves, Lucas, et al.
Published: (2023)
by: Goncalves, Lucas, et al.
Published: (2023)
PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse
by: Vaaras, Einari, et al.
Published: (2024)
by: Vaaras, Einari, et al.
Published: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
by: Jezidžić, Marin, et al.
Published: (2024)
by: Jezidžić, Marin, et al.
Published: (2024)
DeepForestSound: a multi-species automatic detector for passive acoustic monitoring in African tropical forests, a case study in Kibale National Park
by: Dubus, Gabriel, et al.
Published: (2026)
by: Dubus, Gabriel, et al.
Published: (2026)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
by: Pepino, Leonardo, et al.
Published: (2023)
by: Pepino, Leonardo, et al.
Published: (2023)
Cough activity detection for automatic tuberculosis screening
by: van Vüren, Joshua Jansen, et al.
Published: (2026)
by: van Vüren, Joshua Jansen, et al.
Published: (2026)
Investigating self-supervised representations for audio-visual deepfake detection
by: Boldisor, Dragos-Alexandru, et al.
Published: (2025)
by: Boldisor, Dragos-Alexandru, et al.
Published: (2025)
Optimising MFCC parameters for the automatic detection of respiratory diseases
by: Yan, Yuyang, et al.
Published: (2024)
by: Yan, Yuyang, et al.
Published: (2024)
Sentiment analysis in non-fixed length audios using a Fully Convolutional Neural Network
by: García-Ordás, María Teresa, et al.
Published: (2024)
by: García-Ordás, María Teresa, et al.
Published: (2024)
Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
by: Hummel, Hilde I., et al.
Published: (2026)
by: Hummel, Hilde I., et al.
Published: (2026)
Recomposer: Event-roll-guided generative audio editing
by: Ellis, Daniel P. W., et al.
Published: (2025)
by: Ellis, Daniel P. W., et al.
Published: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
by: Yao, Zengwei, et al.
Published: (2023)
by: Yao, Zengwei, et al.
Published: (2023)
Audio-based automatic mating success prediction of giant pandas
by: Yan, WeiRan, et al.
Published: (2019)
by: Yan, WeiRan, et al.
Published: (2019)
Robustifying automatic speech recognition by extracting slowly varying features
by: Pizarro, Matías, et al.
Published: (2021)
by: Pizarro, Matías, et al.
Published: (2021)
Evaluating Interactive 2D Visualization as a Sample Selection Strategy for Biomedical Time-Series Data Annotation
by: Vaaras, Einari, et al.
Published: (2026)
by: Vaaras, Einari, et al.
Published: (2026)
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
by: Siuzdak, Hubert
Published: (2023)
by: Siuzdak, Hubert
Published: (2023)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
by: Robinson, David, et al.
Published: (2024)
by: Robinson, David, et al.
Published: (2024)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
by: Lin, Tsung-En, et al.
Published: (2025)
by: Lin, Tsung-En, et al.
Published: (2025)
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
by: Serrà, Joan, et al.
Published: (2025)
by: Serrà, Joan, et al.
Published: (2025)
Exploring bat song syllable representations in self-supervised audio encoders
by: Kloots, Marianne de Heer, et al.
Published: (2024)
by: Kloots, Marianne de Heer, et al.
Published: (2024)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
by: Fernández-Díaz, Miguel, et al.
Published: (2024)
by: Fernández-Díaz, Miguel, et al.
Published: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
by: Sasindran, Zitha, et al.
Published: (2024)
by: Sasindran, Zitha, et al.
Published: (2024)
Similar Items
-
Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
by: Räsänen, Okko, et al.
Published: (2024) -
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
by: Khorrami, Khazar, et al.
Published: (2021) -
IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
by: Padhya, Dinanath, et al.
Published: (2026) -
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
by: Ferreira, Alexandre R., et al.
Published: (2023) -
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
by: Pahar, Madhurananda, et al.
Published: (2025)