A dataset and model for auditory scene recognition for hearing devices: AHEAD-DS and OpenYAMNet
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Henry, Buchholz, Jörg M., Maclaren, Julian, Carlile, Simon, Lyon, Richard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Identifying Hearing Difficulty Moments in Conversational Audio
von: Collins, Jack, et al.
Veröffentlicht: (2025)
von: Collins, Jack, et al.
Veröffentlicht: (2025)
Controllable joint noise reduction and hearing loss compensation using a differentiable auditory model
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2025)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2025)
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
von: Lachenani, Sidahmed, et al.
Veröffentlicht: (2025)
von: Lachenani, Sidahmed, et al.
Veröffentlicht: (2025)
AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration
von: Zhong, Henry, et al.
Veröffentlicht: (2026)
von: Zhong, Henry, et al.
Veröffentlicht: (2026)
Linear stimulus reconstruction works on the KU Leuven audiovisual, gaze-controlled auditory attention decoding dataset
von: Geirnaert, Simon, et al.
Veröffentlicht: (2024)
von: Geirnaert, Simon, et al.
Veröffentlicht: (2024)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)
Combined assessment of auditory distance perception and externalization
von: Hoppe, Henning, et al.
Veröffentlicht: (2024)
von: Hoppe, Henning, et al.
Veröffentlicht: (2024)
On the relationship between speech and hearing
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
Communication conditions in virtual acoustic scenes in an underground station
von: Hládek, Ľuboš, et al.
Veröffentlicht: (2021)
von: Hládek, Ľuboš, et al.
Veröffentlicht: (2021)
Effects of auditory distance cues and reverberation on spatial perception and listening strategies
von: Missoni, Fulvio, et al.
Veröffentlicht: (2025)
von: Missoni, Fulvio, et al.
Veröffentlicht: (2025)
Loss functions incorporating auditory spatial perception in deep learning -- a review
von: Rafaely, Boaz, et al.
Veröffentlicht: (2025)
von: Rafaely, Boaz, et al.
Veröffentlicht: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Some clues to build a sound analysis relevant to hearing
von: Millot, Laurent
Veröffentlicht: (2024)
von: Millot, Laurent
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Signal processing algorithm effective for sound quality of hearing loss simulators
von: Irino, Toshio, et al.
Veröffentlicht: (2024)
von: Irino, Toshio, et al.
Veröffentlicht: (2024)
How does the teacher rate? Observations from the NeuroPiano dataset
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
Do neonates hear what we measure? Assessing neonatal ward soundscapes at the neonates ears
von: Lam, Bhan, et al.
Veröffentlicht: (2025)
von: Lam, Bhan, et al.
Veröffentlicht: (2025)
Enhancing spatial hearing with cochlear implants: exploring the role of AI, multimodal interaction and perceptual training
von: Picinali, Lorenzo, et al.
Veröffentlicht: (2026)
von: Picinali, Lorenzo, et al.
Veröffentlicht: (2026)
Disentangling peripheral hearing loss from central and cognitive effects on speech intelligibility in older adults
von: Irino, Toshio, et al.
Veröffentlicht: (2025)
von: Irino, Toshio, et al.
Veröffentlicht: (2025)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
Guiding the underwater acoustic target recognition with interpretable contrastive learning
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
Real-time multichannel deep speech enhancement in hearing aids: Comparing monaural and binaural processing in complex acoustic scenarios
von: Westhausen, Nils L., et al.
Veröffentlicht: (2024)
von: Westhausen, Nils L., et al.
Veröffentlicht: (2024)
A contrastive-learning approach for auditory attention detection
von: Bajestan, Seyed Ali Alavi, et al.
Veröffentlicht: (2024)
von: Bajestan, Seyed Ali Alavi, et al.
Veröffentlicht: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Towards interpretable emotion recognition: Identifying key features with machine learning
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
von: Chen, Peijie, et al.
Veröffentlicht: (2025)
von: Chen, Peijie, et al.
Veröffentlicht: (2025)
Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
Multispecies bird sound recognition using a fully convolutional neural network
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions
von: Niu, Shu-Tong, et al.
Veröffentlicht: (2024)
von: Niu, Shu-Tong, et al.
Veröffentlicht: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Fast-Converging Distributed Signal Estimation in Topology-Unconstrained Wireless Acoustic Sensor Networks
von: Didier, Paul, et al.
Veröffentlicht: (2025)
von: Didier, Paul, et al.
Veröffentlicht: (2025)
Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
von: Jiang, Pan-Pan, et al.
Veröffentlicht: (2024)
von: Jiang, Pan-Pan, et al.
Veröffentlicht: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
von: Meng, Ming, et al.
Veröffentlicht: (2025)
von: Meng, Ming, et al.
Veröffentlicht: (2025)
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
von: Lau, Kin Wai, et al.
Veröffentlicht: (2024)
von: Lau, Kin Wai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Identifying Hearing Difficulty Moments in Conversational Audio
von: Collins, Jack, et al.
Veröffentlicht: (2025) -
Controllable joint noise reduction and hearing loss compensation using a differentiable auditory model
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2025) -
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
von: Lachenani, Sidahmed, et al.
Veröffentlicht: (2025) -
AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration
von: Zhong, Henry, et al.
Veröffentlicht: (2026) -
Linear stimulus reconstruction works on the KU Leuven audiovisual, gaze-controlled auditory attention decoding dataset
von: Geirnaert, Simon, et al.
Veröffentlicht: (2024)