Guardado en:
| Autores principales: | Vercammen, Charlotte, Heinrich, Antje, Lesimple, Christophe, Paglialonga, Alessia, Wasmann, Jan-Willem A., Buhl, Mareike |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2505.04728 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Discrimination loss vs. SRT: A model-based approach towards harmonizing speech test interpretations
por: Buhl, Mareike, et al.
Publicado: (2025)
por: Buhl, Mareike, et al.
Publicado: (2025)
Integrating audiological datasets via federated merging of Auditory Profiles
por: Saak, Samira, et al.
Publicado: (2024)
por: Saak, Samira, et al.
Publicado: (2024)
The UmboMic: A PVDF Cantilever Microphone
por: Yeiser, Aaron J., et al.
Publicado: (2023)
por: Yeiser, Aaron J., et al.
Publicado: (2023)
An Implantable Piezofilm Middle Ear Microphone: Performance in Human Cadaveric Temporal Bones
por: Zhang, John Z., et al.
Publicado: (2023)
por: Zhang, John Z., et al.
Publicado: (2023)
On the relevance of acoustic measurements for creating realistic virtual acoustic environments
por: Gündert, Siegfried, et al.
Publicado: (2023)
por: Gündert, Siegfried, et al.
Publicado: (2023)
Standardized Evaluation of Fetal Phonocardiography Processing Methods
por: Müller, Kristóf, et al.
Publicado: (2025)
por: Müller, Kristóf, et al.
Publicado: (2025)
Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
por: Wilkinghoff, Kevin, et al.
Publicado: (2025)
por: Wilkinghoff, Kevin, et al.
Publicado: (2025)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
por: Daeglau, Mareike, et al.
Publicado: (2025)
por: Daeglau, Mareike, et al.
Publicado: (2025)
SALT: Standardized Audio event Label Taxonomy
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)
Timbre Perception, Representation, and its Neuroscientific Exploration: A Comprehensive Review
por: Zhang, Hong, et al.
Publicado: (2024)
por: Zhang, Hong, et al.
Publicado: (2024)
Diff-MST: Differentiable Mixing Style Transfer
por: Vanka, Soumya Sai, et al.
Publicado: (2024)
por: Vanka, Soumya Sai, et al.
Publicado: (2024)
DAC-JAX: A JAX Implementation of the Descript Audio Codec
por: Braun, David
Publicado: (2024)
por: Braun, David
Publicado: (2024)
Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
por: Wilkinghoff, Kevin, et al.
Publicado: (2024)
por: Wilkinghoff, Kevin, et al.
Publicado: (2024)
IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA)
por: Kheir, Yassine El, et al.
Publicado: (2026)
por: Kheir, Yassine El, et al.
Publicado: (2026)
Audio Generation Through Score-Based Generative Modeling: Design Principles and Implementation
por: Zhu, Ge, et al.
Publicado: (2025)
por: Zhu, Ge, et al.
Publicado: (2025)
Implementation and Applications of WakeWords Integrated with Speaker Recognition: A Case Study
por: Filho, Alexandre Costa Ferro, et al.
Publicado: (2024)
por: Filho, Alexandre Costa Ferro, et al.
Publicado: (2024)
Diff-MSTC: A Mixing Style Transfer Prototype for Cubase
por: Vanka, Soumya Sai, et al.
Publicado: (2024)
por: Vanka, Soumya Sai, et al.
Publicado: (2024)
Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
por: Abeßer, Jakob, et al.
Publicado: (2025)
por: Abeßer, Jakob, et al.
Publicado: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
por: Kunešová, Marie, et al.
Publicado: (2025)
por: Kunešová, Marie, et al.
Publicado: (2025)
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
por: Yang, Dong, et al.
Publicado: (2024)
por: Yang, Dong, et al.
Publicado: (2024)
Automatic Music Mixing using a Generative Model of Effect Embeddings
por: Moliner, Eloi, et al.
Publicado: (2025)
por: Moliner, Eloi, et al.
Publicado: (2025)
A Data-Driven Exploration of Elevation Cues in HRTFs: An Explainable AI Perspective Across Multiple Datasets
por: De Rus, Juan Antonio, et al.
Publicado: (2025)
por: De Rus, Juan Antonio, et al.
Publicado: (2025)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
por: Cheng, Yifan, et al.
Publicado: (2025)
por: Cheng, Yifan, et al.
Publicado: (2025)
Streaming Audio Transformers for Online Audio Tagging
por: Dinkel, Heinrich, et al.
Publicado: (2023)
por: Dinkel, Heinrich, et al.
Publicado: (2023)
Scaling up masked audio encoder learning for general audio classification
por: Dinkel, Heinrich, et al.
Publicado: (2024)
por: Dinkel, Heinrich, et al.
Publicado: (2024)
Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
por: Sun, Xingwei, et al.
Publicado: (2025)
por: Sun, Xingwei, et al.
Publicado: (2025)
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
por: Takano, Taisei, et al.
Publicado: (2026)
por: Takano, Taisei, et al.
Publicado: (2026)
Sound Field Translation and Mixed Source Model for Virtual Applications with Perceptual Validation
por: Birnie, Lachlan, et al.
Publicado: (2020)
por: Birnie, Lachlan, et al.
Publicado: (2020)
E2E-AEC: Implementing an end-to-end neural network learning approach for acoustic echo cancellation
por: Jiang, Yiheng, et al.
Publicado: (2026)
por: Jiang, Yiheng, et al.
Publicado: (2026)
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
por: Zhang, Junbo, et al.
Publicado: (2025)
por: Zhang, Junbo, et al.
Publicado: (2025)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
por: Raj, Desh
Publicado: (2024)
por: Raj, Desh
Publicado: (2024)
A New Perspective on Speaker Verification: Joint Modeling with DFSMN and Transformer
por: Wang, Hongyu, et al.
Publicado: (2023)
por: Wang, Hongyu, et al.
Publicado: (2023)
Frequency-Domain Sound Field from the Perspective of Band-Limited Functions
por: Iwami, Takahiro, et al.
Publicado: (2024)
por: Iwami, Takahiro, et al.
Publicado: (2024)
AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
por: Carvalho, Carlos, et al.
Publicado: (2024)
por: Carvalho, Carlos, et al.
Publicado: (2024)
Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
por: Leschanowsky, Anna, et al.
Publicado: (2024)
por: Leschanowsky, Anna, et al.
Publicado: (2024)
Speech Recognition for Analysis of Police Radio Communication
por: Srivastava, Tejes, et al.
Publicado: (2024)
por: Srivastava, Tejes, et al.
Publicado: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
por: Wu, Shih-Lun, et al.
Publicado: (2023)
por: Wu, Shih-Lun, et al.
Publicado: (2023)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
por: Han, Seungu, et al.
Publicado: (2026)
por: Han, Seungu, et al.
Publicado: (2026)
Communication conditions in virtual acoustic scenes in an underground station
por: Hládek, Ľuboš, et al.
Publicado: (2021)
por: Hládek, Ľuboš, et al.
Publicado: (2021)
Ejemplares similares
-
Discrimination loss vs. SRT: A model-based approach towards harmonizing speech test interpretations
por: Buhl, Mareike, et al.
Publicado: (2025) -
Integrating audiological datasets via federated merging of Auditory Profiles
por: Saak, Samira, et al.
Publicado: (2024) -
The UmboMic: A PVDF Cantilever Microphone
por: Yeiser, Aaron J., et al.
Publicado: (2023) -
An Implantable Piezofilm Middle Ear Microphone: Performance in Human Cadaveric Temporal Bones
por: Zhang, John Z., et al.
Publicado: (2023) -
On the relevance of acoustic measurements for creating realistic virtual acoustic environments
por: Gündert, Siegfried, et al.
Publicado: (2023)