Salvato in:
| Autori principali: | Das, Abhijit, Dutta, Sayantan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.06599 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
di: Hinrichs, Reemt, et al.
Pubblicazione: (2025)
di: Hinrichs, Reemt, et al.
Pubblicazione: (2025)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
di: Narain, Jaya, et al.
Pubblicazione: (2025)
di: Narain, Jaya, et al.
Pubblicazione: (2025)
Multi-Source Music Generation with Latent Diffusion
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024)
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024)
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
di: Garcia, Elliot Q C, et al.
Pubblicazione: (2025)
di: Garcia, Elliot Q C, et al.
Pubblicazione: (2025)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
di: Dutta, Soumya, et al.
Pubblicazione: (2024)
di: Dutta, Soumya, et al.
Pubblicazione: (2024)
On the Joint Minimization of Regularization Loss Functions in Deep Variational Bayesian Methods for Attribute-Controlled Symbolic Music Generation
di: Pettenó, Matteo, et al.
Pubblicazione: (2025)
di: Pettenó, Matteo, et al.
Pubblicazione: (2025)
Mathematical Foundations of Polyphonic Music Generation via Structural Inductive Bias
di: Seo, Joonwon
Pubblicazione: (2026)
di: Seo, Joonwon
Pubblicazione: (2026)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
di: Watanabe, Chihiro, et al.
Pubblicazione: (2024)
di: Watanabe, Chihiro, et al.
Pubblicazione: (2024)
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space
di: Rodriguez, Armani, et al.
Pubblicazione: (2024)
di: Rodriguez, Armani, et al.
Pubblicazione: (2024)
Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier
di: Hayta, Berk, et al.
Pubblicazione: (2026)
di: Hayta, Berk, et al.
Pubblicazione: (2026)
Unsupervised Speaker Diarization in Distributed IoT Networks Using Federated Learning
di: Bhuyan, Amit Kumar, et al.
Pubblicazione: (2024)
di: Bhuyan, Amit Kumar, et al.
Pubblicazione: (2024)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
di: Wang, Yiming, et al.
Pubblicazione: (2023)
di: Wang, Yiming, et al.
Pubblicazione: (2023)
Improving the Speaker Anonymization Evaluation's Robustness to Target Speakers with Adversarial Learning
di: Franzreb, Carlos, et al.
Pubblicazione: (2025)
di: Franzreb, Carlos, et al.
Pubblicazione: (2025)
On the Condition Monitoring of Bolted Joints through Acoustic Emission and Deep Transfer Learning: Generalization, Ordinal Loss and Super-Convergence
di: Ramasso, Emmanuel, et al.
Pubblicazione: (2024)
di: Ramasso, Emmanuel, et al.
Pubblicazione: (2024)
Voice Disorder Analysis: a Transformer-based Approach
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses
di: Xu, Shaoheng, et al.
Pubblicazione: (2026)
di: Xu, Shaoheng, et al.
Pubblicazione: (2026)
Do Foundational Audio Encoders Understand Music Structure?
di: Toyama, Keisuke, et al.
Pubblicazione: (2025)
di: Toyama, Keisuke, et al.
Pubblicazione: (2025)
Parameter-Efficient Transfer Learning for Music Foundation Models
di: Ding, Yiwei, et al.
Pubblicazione: (2024)
di: Ding, Yiwei, et al.
Pubblicazione: (2024)
Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
di: Gerazov, Branislav, et al.
Pubblicazione: (2025)
di: Gerazov, Branislav, et al.
Pubblicazione: (2025)
Safeguarding Privacy in Edge Speech Understanding with Tiny Foundation Models
di: Benazir, Afsara, et al.
Pubblicazione: (2025)
di: Benazir, Afsara, et al.
Pubblicazione: (2025)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
di: Sarkar, Eklavya, et al.
Pubblicazione: (2024)
di: Sarkar, Eklavya, et al.
Pubblicazione: (2024)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
An Independence-promoting Loss for Music Generation with Language Models
di: Lemercier, Jean-Marie, et al.
Pubblicazione: (2024)
di: Lemercier, Jean-Marie, et al.
Pubblicazione: (2024)
Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation
di: Nie, Jingping, et al.
Pubblicazione: (2025)
di: Nie, Jingping, et al.
Pubblicazione: (2025)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
di: Zhao, Shengkui, et al.
Pubblicazione: (2021)
di: Zhao, Shengkui, et al.
Pubblicazione: (2021)
Automatic Identification of Samples in Hip-Hop Music via Multi-Loss Training and an Artificial Dataset
di: Cheston, Huw, et al.
Pubblicazione: (2025)
di: Cheston, Huw, et al.
Pubblicazione: (2025)
Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask Estimation
di: Kienegger, Jakob, et al.
Pubblicazione: (2024)
di: Kienegger, Jakob, et al.
Pubblicazione: (2024)
Anticipatory Music Transformer
di: Thickstun, John, et al.
Pubblicazione: (2023)
di: Thickstun, John, et al.
Pubblicazione: (2023)
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
Unsupervised CP-UNet Framework for Denoising DAS Data with Decay Noise
di: Huang, Tianye, et al.
Pubblicazione: (2025)
di: Huang, Tianye, et al.
Pubblicazione: (2025)
Patient-Aware Feature Alignment for Robust Lung Sound Classification:Cohesion-Separation and Global Alignment Losses
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2025)
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2025)
BiFormer3D: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer
di: Xu, Shaoheng, et al.
Pubblicazione: (2026)
di: Xu, Shaoheng, et al.
Pubblicazione: (2026)
A Power-Weighted Noncentral Complex Gaussian Distribution
di: Nakashika, Toru
Pubblicazione: (2026)
di: Nakashika, Toru
Pubblicazione: (2026)
USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
di: Zhao, Guanlong, et al.
Pubblicazione: (2023)
di: Zhao, Guanlong, et al.
Pubblicazione: (2023)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
di: Ashvin, Aditya, et al.
Pubblicazione: (2024)
di: Ashvin, Aditya, et al.
Pubblicazione: (2024)
E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
di: Dong, Jiaheng, et al.
Pubblicazione: (2025)
di: Dong, Jiaheng, et al.
Pubblicazione: (2025)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
MK-SGC-SC: Multiple Kernel Guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization
di: Raghav, Nikhil, et al.
Pubblicazione: (2026)
di: Raghav, Nikhil, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting
di: Xiao, Yang, et al.
Pubblicazione: (2025) -
Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
di: Hinrichs, Reemt, et al.
Pubblicazione: (2025) -
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
di: Narain, Jaya, et al.
Pubblicazione: (2025) -
Multi-Source Music Generation with Latent Diffusion
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024) -
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
di: Garcia, Elliot Q C, et al.
Pubblicazione: (2025)