Is Attention always needed? A Case Study on Language Identification from Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Mandal, Atanu, Pal, Santanu, Dutta, Indranil, Bhattacharya, Mahidas, Naskar, Sudip Kumar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2021
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
di: Wills, Simone, et al.
Pubblicazione: (2023)
di: Wills, Simone, et al.
Pubblicazione: (2023)
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
di: Goswami, Mandip
Pubblicazione: (2026)
di: Goswami, Mandip
Pubblicazione: (2026)
A Study on Speech Assessment with Visual Cues
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
di: Zhang, Wangyou, et al.
Pubblicazione: (2025)
di: Zhang, Wangyou, et al.
Pubblicazione: (2025)
Semantic Communications for Speech Recognition
di: Weng, Zhenzi, et al.
Pubblicazione: (2021)
di: Weng, Zhenzi, et al.
Pubblicazione: (2021)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
di: Wilroth, Johanna, et al.
Pubblicazione: (2026)
di: Wilroth, Johanna, et al.
Pubblicazione: (2026)
Toward Universal Speech Enhancement for Diverse Input Conditions
di: Zhang, Wangyou, et al.
Pubblicazione: (2023)
di: Zhang, Wangyou, et al.
Pubblicazione: (2023)
Speech dereverberation constrained on room impulse response characteristics
di: Bahrman, Louis, et al.
Pubblicazione: (2024)
di: Bahrman, Louis, et al.
Pubblicazione: (2024)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
di: Joysingh, S. Johanan, et al.
Pubblicazione: (2024)
di: Joysingh, S. Johanan, et al.
Pubblicazione: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
di: Gupta, Kishan, et al.
Pubblicazione: (2024)
di: Gupta, Kishan, et al.
Pubblicazione: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
di: Yu, Chin-Yun, et al.
Pubblicazione: (2022)
di: Yu, Chin-Yun, et al.
Pubblicazione: (2022)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
di: Rahimi, Akam, et al.
Pubblicazione: (2025)
di: Rahimi, Akam, et al.
Pubblicazione: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
di: Serre, Thomas, et al.
Pubblicazione: (2026)
di: Serre, Thomas, et al.
Pubblicazione: (2026)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
di: Kwon, Younghoo, et al.
Pubblicazione: (2024)
di: Kwon, Younghoo, et al.
Pubblicazione: (2024)
Binaural Selective Attention Model for Target Speaker Extraction
di: Meng, Hanyu, et al.
Pubblicazione: (2024)
di: Meng, Hanyu, et al.
Pubblicazione: (2024)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
di: Yuan, Ze, et al.
Pubblicazione: (2024)
di: Yuan, Ze, et al.
Pubblicazione: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
di: Chen, Xiaodan, et al.
Pubblicazione: (2025)
di: Chen, Xiaodan, et al.
Pubblicazione: (2025)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
Speech-preserving active noise control: a deep learning approach in reverberant environments
di: Dai, Shuning
Pubblicazione: (2026)
di: Dai, Shuning
Pubblicazione: (2026)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
di: Yang, Yujie, et al.
Pubblicazione: (2025)
di: Yang, Yujie, et al.
Pubblicazione: (2025)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2023)
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2023)
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
di: Zhu, Haolin, et al.
Pubblicazione: (2024)
di: Zhu, Haolin, et al.
Pubblicazione: (2024)
AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
di: Gong, Xun, et al.
Pubblicazione: (2025)
di: Gong, Xun, et al.
Pubblicazione: (2025)
Detecting Post-Stroke Aphasia Via Brain Responses to Speech in a Deep Learning Framework
di: De Clercq, Pieter, et al.
Pubblicazione: (2024)
di: De Clercq, Pieter, et al.
Pubblicazione: (2024)
Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
di: Tang, Chenyu, et al.
Pubblicazione: (2023)
di: Tang, Chenyu, et al.
Pubblicazione: (2023)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
di: Yan, Haoyin, et al.
Pubblicazione: (2024)
di: Yan, Haoyin, et al.
Pubblicazione: (2024)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
di: Hao, Xiang, et al.
Pubblicazione: (2020)
di: Hao, Xiang, et al.
Pubblicazione: (2020)
Direct Speech-to-Speech Neural Machine Translation: A Survey
di: Gupta, Mahendra, et al.
Pubblicazione: (2024)
di: Gupta, Mahendra, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
di: Mandal, Atanu, et al.
Pubblicazione: (2024) -
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
di: Wills, Simone, et al.
Pubblicazione: (2023) -
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
di: Goswami, Mandip
Pubblicazione: (2026) -
A Study on Speech Assessment with Visual Cues
di: Ahmed, Shafique, et al.
Pubblicazione: (2025) -
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)