Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Das, Abhijit, Dutta, Sayantan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
von: Hinrichs, Reemt, et al.
Veröffentlicht: (2025)
von: Hinrichs, Reemt, et al.
Veröffentlicht: (2025)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
von: Garcia, Elliot Q C, et al.
Veröffentlicht: (2025)
von: Garcia, Elliot Q C, et al.
Veröffentlicht: (2025)
Multi-Source Music Generation with Latent Diffusion
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space
von: Rodriguez, Armani, et al.
Veröffentlicht: (2024)
von: Rodriguez, Armani, et al.
Veröffentlicht: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
Mathematical Foundations of Polyphonic Music Generation via Structural Inductive Bias
von: Seo, Joonwon
Veröffentlicht: (2026)
von: Seo, Joonwon
Veröffentlicht: (2026)
On the Joint Minimization of Regularization Loss Functions in Deep Variational Bayesian Methods for Attribute-Controlled Symbolic Music Generation
von: Pettenó, Matteo, et al.
Veröffentlicht: (2025)
von: Pettenó, Matteo, et al.
Veröffentlicht: (2025)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier
von: Hayta, Berk, et al.
Veröffentlicht: (2026)
von: Hayta, Berk, et al.
Veröffentlicht: (2026)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Improving the Speaker Anonymization Evaluation's Robustness to Target Speakers with Adversarial Learning
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)
On the Condition Monitoring of Bolted Joints through Acoustic Emission and Deep Transfer Learning: Generalization, Ordinal Loss and Super-Convergence
von: Ramasso, Emmanuel, et al.
Veröffentlicht: (2024)
von: Ramasso, Emmanuel, et al.
Veröffentlicht: (2024)
Voice Disorder Analysis: a Transformer-based Approach
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
Unsupervised Speaker Diarization in Distributed IoT Networks Using Federated Learning
von: Bhuyan, Amit Kumar, et al.
Veröffentlicht: (2024)
von: Bhuyan, Amit Kumar, et al.
Veröffentlicht: (2024)
RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses
von: Xu, Shaoheng, et al.
Veröffentlicht: (2026)
von: Xu, Shaoheng, et al.
Veröffentlicht: (2026)
Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
von: Gerazov, Branislav, et al.
Veröffentlicht: (2025)
von: Gerazov, Branislav, et al.
Veröffentlicht: (2025)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
von: Wang, Yiming, et al.
Veröffentlicht: (2023)
von: Wang, Yiming, et al.
Veröffentlicht: (2023)
Do Foundational Audio Encoders Understand Music Structure?
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
Parameter-Efficient Transfer Learning for Music Foundation Models
von: Ding, Yiwei, et al.
Veröffentlicht: (2024)
von: Ding, Yiwei, et al.
Veröffentlicht: (2024)
BiFormer3D: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer
von: Xu, Shaoheng, et al.
Veröffentlicht: (2026)
von: Xu, Shaoheng, et al.
Veröffentlicht: (2026)
Safeguarding Privacy in Edge Speech Understanding with Tiny Foundation Models
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation
von: Nie, Jingping, et al.
Veröffentlicht: (2025)
von: Nie, Jingping, et al.
Veröffentlicht: (2025)
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
von: Richter, Julius, et al.
Veröffentlicht: (2026)
von: Richter, Julius, et al.
Veröffentlicht: (2026)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
Automatic Identification of Samples in Hip-Hop Music via Multi-Loss Training and an Artificial Dataset
von: Cheston, Huw, et al.
Veröffentlicht: (2025)
von: Cheston, Huw, et al.
Veröffentlicht: (2025)
Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask Estimation
von: Kienegger, Jakob, et al.
Veröffentlicht: (2024)
von: Kienegger, Jakob, et al.
Veröffentlicht: (2024)
Anticipatory Music Transformer
von: Thickstun, John, et al.
Veröffentlicht: (2023)
von: Thickstun, John, et al.
Veröffentlicht: (2023)
Patient-Aware Feature Alignment for Robust Lung Sound Classification:Cohesion-Separation and Global Alignment Losses
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
von: Singh, Parampreet, et al.
Veröffentlicht: (2026)
von: Singh, Parampreet, et al.
Veröffentlicht: (2026)
PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
von: Chernyak, Bronya Roni, et al.
Veröffentlicht: (2025)
von: Chernyak, Bronya Roni, et al.
Veröffentlicht: (2025)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
von: Lovelace, Justin, et al.
Veröffentlicht: (2025)
von: Lovelace, Justin, et al.
Veröffentlicht: (2025)
USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
von: Zhao, Guanlong, et al.
Veröffentlicht: (2023)
von: Zhao, Guanlong, et al.
Veröffentlicht: (2023)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
von: Dong, Jiaheng, et al.
Veröffentlicht: (2025)
von: Dong, Jiaheng, et al.
Veröffentlicht: (2025)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting
von: Xiao, Yang, et al.
Veröffentlicht: (2025) -
Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
von: Hinrichs, Reemt, et al.
Veröffentlicht: (2025) -
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
von: Narain, Jaya, et al.
Veröffentlicht: (2025) -
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
von: Garcia, Elliot Q C, et al.
Veröffentlicht: (2025) -
Multi-Source Music Generation with Latent Diffusion
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)