Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Amartyaveer, Kadambi, Murali, Sharma, Chandra Mohan, Mondal, Anupam, Ghosh, Prasanta Kumar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentation
di: Tholan, Masoud Thajudeen, et al.
Pubblicazione: (2025)
di: Tholan, Masoud Thajudeen, et al.
Pubblicazione: (2025)
A Generalized Weighted Overlap-Add (WOLA) Filter Bank for Improved Subband System Identification
di: Sharma, Mohit, et al.
Pubblicazione: (2025)
di: Sharma, Mohit, et al.
Pubblicazione: (2025)
Incremental Averaging Method to Improve Graph-Based Time-Difference-of-Arrival Estimation
di: Brümann, Klaus, et al.
Pubblicazione: (2025)
di: Brümann, Klaus, et al.
Pubblicazione: (2025)
Modeling and Link Budget Feasibility Analysis of Secure LoRa-Based Peer-to-Peer Communication for Short-Range Tactical Networks
di: Agrawal, Ayush Kumar, et al.
Pubblicazione: (2026)
di: Agrawal, Ayush Kumar, et al.
Pubblicazione: (2026)
Advanced Signal Analysis in Detecting Replay Attacks for Automatic Speaker Verification Systems
di: Kuang, Lee Shih
Pubblicazione: (2024)
di: Kuang, Lee Shih
Pubblicazione: (2024)
Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2026)
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2026)
Heart Murmur and Abnormal PCG Detection via Wavelet Scattering Transform & a 1D-CNN
di: Patwa, Ahmed, et al.
Pubblicazione: (2023)
di: Patwa, Ahmed, et al.
Pubblicazione: (2023)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
di: Xie, Yuying, et al.
Pubblicazione: (2024)
di: Xie, Yuying, et al.
Pubblicazione: (2024)
A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays
di: Miotello, Federico, et al.
Pubblicazione: (2024)
di: Miotello, Federico, et al.
Pubblicazione: (2024)
Ultra Low Complexity Deep Learning Based Noise Suppression
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2023)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2023)
Automatic Voice Classification Of Autistic Subjects
di: Vacca, Jessica, et al.
Pubblicazione: (2024)
di: Vacca, Jessica, et al.
Pubblicazione: (2024)
Resource-Efficient Separation Transformer
di: Della Libera, Luca, et al.
Pubblicazione: (2022)
di: Della Libera, Luca, et al.
Pubblicazione: (2022)
Ambisonics Encoder for Wearable Array with Improved Binaural Reproduction
di: Gayer, Yhonatan, et al.
Pubblicazione: (2025)
di: Gayer, Yhonatan, et al.
Pubblicazione: (2025)
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers
di: Mang, Loredana Daria, et al.
Pubblicazione: (2024)
di: Mang, Loredana Daria, et al.
Pubblicazione: (2024)
Detection of manatee vocalisations using the Audio Spectrogram Transformer
di: Schiappacasse, Stefano, et al.
Pubblicazione: (2024)
di: Schiappacasse, Stefano, et al.
Pubblicazione: (2024)
Musical Score Following using Statistical Inference
di: Cowley, Josephine
Pubblicazione: (2025)
di: Cowley, Josephine
Pubblicazione: (2025)
AutoMashup: Automatic Music Mashups Creation
di: Delabaere, Marine, et al.
Pubblicazione: (2025)
di: Delabaere, Marine, et al.
Pubblicazione: (2025)
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Improving Machine Hearing on Limited Data Sets
di: Harar, Pavol, et al.
Pubblicazione: (2019)
di: Harar, Pavol, et al.
Pubblicazione: (2019)
Improved in-car sound pick-up using multichannel Wiener filter
di: Khalid, Juhi, et al.
Pubblicazione: (2025)
di: Khalid, Juhi, et al.
Pubblicazione: (2025)
Reduce Computational Complexity for Continuous Wavelet Transform in Acoustic Recognition Using Hop Size
di: Phan, Dang Thoai
Pubblicazione: (2024)
di: Phan, Dang Thoai
Pubblicazione: (2024)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
di: Baoueb, Teysir, et al.
Pubblicazione: (2025)
di: Baoueb, Teysir, et al.
Pubblicazione: (2025)
On the Invariance of Cross-Correlation Peak Positions Under Monotonic Signal Transformations, with Application to Fast Time Difference Estimation
di: Ueno, Natsuki, et al.
Pubblicazione: (2025)
di: Ueno, Natsuki, et al.
Pubblicazione: (2025)
ILD-VIT: A Unified Vision Transformer Architecture for Detection of Interstitial Lung Disease from Respiratory Sounds
di: Hota, Soubhagya Ranjan, et al.
Pubblicazione: (2025)
di: Hota, Soubhagya Ranjan, et al.
Pubblicazione: (2025)
A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
di: Damiano, Stefano, et al.
Pubblicazione: (2024)
di: Damiano, Stefano, et al.
Pubblicazione: (2024)
Note-Level Singing Melody Transcription for Time-Aligned Musical Score Generation
di: Kim, Leekyung, et al.
Pubblicazione: (2025)
di: Kim, Leekyung, et al.
Pubblicazione: (2025)
Learnable Adaptive Time-Frequency Representation via Differentiable Short-Time Fourier Transform
di: Leiber, Maxime, et al.
Pubblicazione: (2025)
di: Leiber, Maxime, et al.
Pubblicazione: (2025)
Speech-Based Prioritization for Schizophrenia Intervention
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
Predictive Directional Selective Fixed-Filter Active Noise Control for Moving Sources via a Convolutional Recurrent Neural Network
di: Wang, Boxiang, et al.
Pubblicazione: (2026)
di: Wang, Boxiang, et al.
Pubblicazione: (2026)
FunnelNet: An End-to-End Deep Learning Framework to Monitor Digital Heart Murmur in Real-Time
di: Jobayer, Md, et al.
Pubblicazione: (2024)
di: Jobayer, Md, et al.
Pubblicazione: (2024)
Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
di: Bhattacharjee, Sankha Subhra, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Sankha Subhra, et al.
Pubblicazione: (2024)
Blind Capon Beamformer Based on Independent Component Extraction: Single-Parameter Algorithm,
di: Koldovský, Zbyněk, et al.
Pubblicazione: (2025)
di: Koldovský, Zbyněk, et al.
Pubblicazione: (2025)
Multi-Source Position and Direction-of-Arrival Estimation Based on Euclidean Distance Matrices
di: Brümann, Klaus, et al.
Pubblicazione: (2025)
di: Brümann, Klaus, et al.
Pubblicazione: (2025)
Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
di: Liu, Weixin, et al.
Pubblicazione: (2026)
di: Liu, Weixin, et al.
Pubblicazione: (2026)
Transferable Selective Virtual Sensing Active Noise Control Technique Based on Metric Learning
di: Wang, Boxiang, et al.
Pubblicazione: (2024)
di: Wang, Boxiang, et al.
Pubblicazione: (2024)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
di: Pešán, Jan, et al.
Pubblicazione: (2024)
di: Pešán, Jan, et al.
Pubblicazione: (2024)
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
di: Serbest, Sanberk, et al.
Pubblicazione: (2025)
di: Serbest, Sanberk, et al.
Pubblicazione: (2025)
Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality
di: Marinoni, Christian, et al.
Pubblicazione: (2024)
di: Marinoni, Christian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentation
di: Tholan, Masoud Thajudeen, et al.
Pubblicazione: (2025) -
A Generalized Weighted Overlap-Add (WOLA) Filter Bank for Improved Subband System Identification
di: Sharma, Mohit, et al.
Pubblicazione: (2025) -
Incremental Averaging Method to Improve Graph-Based Time-Difference-of-Arrival Estimation
di: Brümann, Klaus, et al.
Pubblicazione: (2025) -
Modeling and Link Budget Feasibility Analysis of Secure LoRa-Based Peer-to-Peer Communication for Short-Range Tactical Networks
di: Agrawal, Ayush Kumar, et al.
Pubblicazione: (2026) -
Advanced Signal Analysis in Detecting Replay Attacks for Automatic Speaker Verification Systems
di: Kuang, Lee Shih
Pubblicazione: (2024)