A Data-Centric Framework for Machine Listening Projects: Addressing Large-Scale Data Acquisition and Labeling through Active Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Naranjo-Alcazar, Javier, Grau-Haro, Jordi, Ribes-Serrano, Ruben, Zuccarello, Pedro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
by: Grau-Haro, Jordi, et al.
Published: (2025)
by: Grau-Haro, Jordi, et al.
Published: (2025)
Female mosquito detection by means of AI techniques inside release containers in the context of a Sterile Insect Technique program
by: Naranjo-Alcazar, Javier, et al.
Published: (2023)
by: Naranjo-Alcazar, Javier, et al.
Published: (2023)
Spike Encoding for Environmental Sound: A Comparative Benchmark
by: Larroza, Andres, et al.
Published: (2025)
by: Larroza, Andres, et al.
Published: (2025)
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
by: Ghosh, Bishal, et al.
Published: (2024)
by: Ghosh, Bishal, et al.
Published: (2024)
RF-GML: Reference-Free Generative Machine Listener
by: Biswas, Arijit, et al.
Published: (2024)
by: Biswas, Arijit, et al.
Published: (2024)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
by: Chung, Soo-Whan, et al.
Published: (2025)
by: Chung, Soo-Whan, et al.
Published: (2025)
Listen, Think, and Understand
by: Gong, Yuan, et al.
Published: (2023)
by: Gong, Yuan, et al.
Published: (2023)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
by: Hu, Cheng-Hung, et al.
Published: (2025)
by: Hu, Cheng-Hung, et al.
Published: (2025)
DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
by: Cornell, Samuele, et al.
Published: (2024)
by: Cornell, Samuele, et al.
Published: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
by: Jeong, Jihoon, et al.
Published: (2026)
by: Jeong, Jihoon, et al.
Published: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
by: Gelderblom, Femke B., et al.
Published: (2024)
by: Gelderblom, Femke B., et al.
Published: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
by: Shen, Pengjie, et al.
Published: (2025)
by: Shen, Pengjie, et al.
Published: (2025)
Less is More: Data Curation Matters in Scaling Speech Enhancement
by: Li, Chenda, et al.
Published: (2025)
by: Li, Chenda, et al.
Published: (2025)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
by: Raj, Desh
Published: (2024)
by: Raj, Desh
Published: (2024)
Requirements for Mass Adoption of Assistive Listening Technology by the General Public
by: Kaufmann, Thomas B., et al.
Published: (2023)
by: Kaufmann, Thomas B., et al.
Published: (2023)
DIFFA: Large Language Diffusion Models Can Listen and Understand
by: Zhou, Jiaming, et al.
Published: (2025)
by: Zhou, Jiaming, et al.
Published: (2025)
Reproducing the Acoustic Velocity Vectors in a Circular Listening Area
by: Wang, Jiarui, et al.
Published: (2024)
by: Wang, Jiarui, et al.
Published: (2024)
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
by: Fuglsig, Andreas J., et al.
Published: (2023)
by: Fuglsig, Andreas J., et al.
Published: (2023)
Listening broadband physical model for microphones: a first step
by: Millot, Laurent, et al.
Published: (2024)
by: Millot, Laurent, et al.
Published: (2024)
Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
by: Lehmkuhl, Jonathan, et al.
Published: (2025)
by: Lehmkuhl, Jonathan, et al.
Published: (2025)
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
by: Choi, Woosung, et al.
Published: (2025)
by: Choi, Woosung, et al.
Published: (2025)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
by: Suda, Hitoshi, et al.
Published: (2024)
by: Suda, Hitoshi, et al.
Published: (2024)
Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
by: Shen, Yu-Han, et al.
Published: (2018)
by: Shen, Yu-Han, et al.
Published: (2018)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
by: Yamamoto, Katsuhiko, et al.
Published: (2025)
by: Yamamoto, Katsuhiko, et al.
Published: (2025)
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
by: Ho, Tuan Vu, et al.
Published: (2024)
by: Ho, Tuan Vu, et al.
Published: (2024)
An Exploratory Study of Multimodal Physiological Data in Jazz Improvisation Using Basic Machine Learning Techniques
by: Zhang, Yawen
Published: (2024)
by: Zhang, Yawen
Published: (2024)
Oral Tradition-Encoded NanyinHGNN: Integrating Nanyin Music Preservation and Generation through a Pipa-Centric Dataset
by: Xiahou, Jianbing, et al.
Published: (2025)
by: Xiahou, Jianbing, et al.
Published: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
by: Chen, Zhengyang, et al.
Published: (2024)
by: Chen, Zhengyang, et al.
Published: (2024)
Water Flow Detection Device Based on Sound Data Analysis and Machine Learning to Detect Water Leakage
by: Pourmehrani, Hossein, et al.
Published: (2025)
by: Pourmehrani, Hossein, et al.
Published: (2025)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
by: Borodin, Kirill, et al.
Published: (2025)
by: Borodin, Kirill, et al.
Published: (2025)
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
by: Kawamura, Takao, et al.
Published: (2026)
by: Kawamura, Takao, et al.
Published: (2026)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
by: Wang, Chunhui, et al.
Published: (2024)
by: Wang, Chunhui, et al.
Published: (2024)
Unraveling Complex Data Diversity in Underwater Acoustic Target Recognition through Convolution-based Mixture of Experts
by: Xie, Yuan, et al.
Published: (2024)
by: Xie, Yuan, et al.
Published: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence
by: Masiero, Bruno S., et al.
Published: (2024)
by: Masiero, Bruno S., et al.
Published: (2024)
Spatial Analysis and Synthesis Methods: Subjective and Objective Evaluations Using Various Microphone Arrays in the Auralization of a Critical Listening Room
by: Pawlak, Alan, et al.
Published: (2024)
by: Pawlak, Alan, et al.
Published: (2024)
AdaProj: Adaptively Scaled Angular Margin Subspace Projections for Anomalous Sound Detection with Auxiliary Classification Tasks
by: Wilkinghoff, Kevin
Published: (2024)
by: Wilkinghoff, Kevin
Published: (2024)
A Multi-loudspeaker Binaural Room Impulse Response Dataset with High-Resolution Translational and Rotational Head Coordinates in a Listening Room
by: Qiao, Yue, et al.
Published: (2024)
by: Qiao, Yue, et al.
Published: (2024)
Similar Items
-
Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
by: Grau-Haro, Jordi, et al.
Published: (2025) -
Female mosquito detection by means of AI techniques inside release containers in the context of a Sterile Insect Technique program
by: Naranjo-Alcazar, Javier, et al.
Published: (2023) -
Spike Encoding for Environmental Sound: A Comparative Benchmark
by: Larroza, Andres, et al.
Published: (2025) -
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
by: Ghosh, Bishal, et al.
Published: (2024) -
RF-GML: Reference-Free Generative Machine Listener
by: Biswas, Arijit, et al.
Published: (2024)