Saved in:
| Main Authors: | Cabansag, Ian Jacob, Ntegeka, Paul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.01391 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prediction of Spotify Chart Success Using Audio and Streaming Features
by: Cabansag, Ian Jacob, et al.
Published: (2025)
by: Cabansag, Ian Jacob, et al.
Published: (2025)
Rethinking Non-Negative Matrix Factorization with Implicit Neural Representations
by: Subramani, Krishna, et al.
Published: (2024)
by: Subramani, Krishna, et al.
Published: (2024)
Bayesian Learning for Deep Neural Network Adaptation
by: Xie, Xurong, et al.
Published: (2020)
by: Xie, Xurong, et al.
Published: (2020)
Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
by: Bauer, Waldemar, et al.
Published: (2024)
by: Bauer, Waldemar, et al.
Published: (2024)
Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
by: Kienegger, Jakob, et al.
Published: (2026)
by: Kienegger, Jakob, et al.
Published: (2026)
RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
by: Bitterman, Jacob, et al.
Published: (2024)
by: Bitterman, Jacob, et al.
Published: (2024)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
by: Primus, Paul, et al.
Published: (2024)
by: Primus, Paul, et al.
Published: (2024)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
by: Kim, Hyeongju, et al.
Published: (2025)
by: Kim, Hyeongju, et al.
Published: (2025)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
by: Primus, Paul, et al.
Published: (2025)
by: Primus, Paul, et al.
Published: (2025)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
by: Primus, Paul, et al.
Published: (2024)
by: Primus, Paul, et al.
Published: (2024)
Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
by: Park, Hyun Jin, et al.
Published: (2024)
by: Park, Hyun Jin, et al.
Published: (2024)
Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
by: Park, Hyun Jin, et al.
Published: (2024)
by: Park, Hyun Jin, et al.
Published: (2024)
Bayesian Low-Rank Factorization for Robust Model Adaptation
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
by: Tal, Or, et al.
Published: (2025)
by: Tal, Or, et al.
Published: (2025)
Bayesian Restoration of Audio Degraded by Low-Frequency Pulses Modeled via Gaussian Process
by: de Carvalho, Hugo Tremonte, et al.
Published: (2020)
by: de Carvalho, Hugo Tremonte, et al.
Published: (2020)
Fine-grained Soundscape Control for Augmented Hearing
by: Oh, Seunghyun, et al.
Published: (2026)
by: Oh, Seunghyun, et al.
Published: (2026)
EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
by: Cerovaz, Luca, et al.
Published: (2026)
by: Cerovaz, Luca, et al.
Published: (2026)
Test-Time Adaptation for Speech Emotion Recognition
by: Dong, Jiaheng, et al.
Published: (2026)
by: Dong, Jiaheng, et al.
Published: (2026)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
by: Sun, Esther, et al.
Published: (2026)
by: Sun, Esther, et al.
Published: (2026)
Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing
by: Hizabri, Marouane El, et al.
Published: (2026)
by: Hizabri, Marouane El, et al.
Published: (2026)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
by: Pahwa, Ramit, et al.
Published: (2026)
by: Pahwa, Ramit, et al.
Published: (2026)
LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classification
by: Jazaeri, Niloofar, et al.
Published: (2026)
by: Jazaeri, Niloofar, et al.
Published: (2026)
Self-Supervised Learning for Speaker Recognition: A study and review
by: Lepage, Theo, et al.
Published: (2026)
by: Lepage, Theo, et al.
Published: (2026)
Evaluating Disentangled Representations for Controllable Music Generation
by: Ibáñez-Martínez, Laura, et al.
Published: (2026)
by: Ibáñez-Martínez, Laura, et al.
Published: (2026)
Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
by: Alexey, Protopopov
Published: (2026)
by: Alexey, Protopopov
Published: (2026)
MK-SGC-SC: Multiple Kernel Guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization
by: Raghav, Nikhil, et al.
Published: (2026)
by: Raghav, Nikhil, et al.
Published: (2026)
SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
by: Yemini, Yochai, et al.
Published: (2026)
by: Yemini, Yochai, et al.
Published: (2026)
Context-aware child-directed speech detection from long-form recordings
by: Charlot, Théo, et al.
Published: (2026)
by: Charlot, Théo, et al.
Published: (2026)
Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech
by: Samanta, Himadri S
Published: (2026)
by: Samanta, Himadri S
Published: (2026)
RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity
by: Bertolino, Gaia A., et al.
Published: (2026)
by: Bertolino, Gaia A., et al.
Published: (2026)
Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
by: Nallaguntla, Vamshi, et al.
Published: (2026)
by: Nallaguntla, Vamshi, et al.
Published: (2026)
From Diet to Free Lunch: Estimating Auxiliary Signal Properties using Dynamic Pruning Masks in Speech Enhancement Networks
by: Miccini, Riccardo, et al.
Published: (2026)
by: Miccini, Riccardo, et al.
Published: (2026)
Multi-Channel Replay Speech Detection using Acoustic Maps
by: Neri, Michael, et al.
Published: (2026)
by: Neri, Michael, et al.
Published: (2026)
Toward Faithful Explanations in Acoustic Anomaly Detection
by: Elrashid, Maab, et al.
Published: (2026)
by: Elrashid, Maab, et al.
Published: (2026)
Frame-Level Internal Tool Use for Temporal Grounding in Audio LMs
by: An, Joesph, et al.
Published: (2026)
by: An, Joesph, et al.
Published: (2026)
SCRAPL: Scattering Transform with Random Paths for Machine Learning
by: Mitcheltree, Christopher, et al.
Published: (2026)
by: Mitcheltree, Christopher, et al.
Published: (2026)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
by: Menon, Aditya Srinivas, et al.
Published: (2026)
by: Menon, Aditya Srinivas, et al.
Published: (2026)
Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations
by: Doerfler, Robin, et al.
Published: (2026)
by: Doerfler, Robin, et al.
Published: (2026)
Similar Items
-
Prediction of Spotify Chart Success Using Audio and Streaming Features
by: Cabansag, Ian Jacob, et al.
Published: (2025) -
Rethinking Non-Negative Matrix Factorization with Implicit Neural Representations
by: Subramani, Krishna, et al.
Published: (2024) -
Bayesian Learning for Deep Neural Network Adaptation
by: Xie, Xurong, et al.
Published: (2020) -
Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
by: Bauer, Waldemar, et al.
Published: (2024) -
Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
by: Kienegger, Jakob, et al.
Published: (2026)