A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tamiti, Tarikul Islam, Joshi, Biraj, Hasan, Rida, Hasan, Rashedul, Athay, Taieba, Mamun, Nursad, Barua, Anomadarshi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WaLi: Can Pressure Sensors in HVAC Systems Capture Human Speech?
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
CIS-BWE: Chaos-Informed Speech Bandwidth Extension
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
SUBARU: A Practical Approach to Power Saving in Hearables Using SUB-Nyquist Audio Resolution Upsampling
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
von: Hasan, Rashedul, et al.
Veröffentlicht: (2025)
von: Hasan, Rashedul, et al.
Veröffentlicht: (2025)
NLDSI-BWE: Non Linear Dynamical Systems-Inspired Multi Resolution Discriminators for Speech Bandwidth Extension
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network
von: Avro, Shamin Bin Habib, et al.
Veröffentlicht: (2025)
von: Avro, Shamin Bin Habib, et al.
Veröffentlicht: (2025)
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
von: Prome, Ruhina Tabasshum, et al.
Veröffentlicht: (2025)
von: Prome, Ruhina Tabasshum, et al.
Veröffentlicht: (2025)
A Fly on the Wall -- Exploiting Acoustic Side-Channels in Differential Pressure Sensors
von: Achamyeleh, Yonatan Gizachew, et al.
Veröffentlicht: (2024)
von: Achamyeleh, Yonatan Gizachew, et al.
Veröffentlicht: (2024)
DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network
von: Mamun, Nursadul, et al.
Veröffentlicht: (2026)
von: Mamun, Nursadul, et al.
Veröffentlicht: (2026)
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025)
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025)
A Self-Attention-Driven Deep Denoiser Model for Real Time Lung Sound Denoising in Noisy Environments
von: Shuvo, Samiul Based, et al.
Veröffentlicht: (2024)
von: Shuvo, Samiul Based, et al.
Veröffentlicht: (2024)
Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
von: N, Rishith Sadashiv T, et al.
Veröffentlicht: (2025)
von: N, Rishith Sadashiv T, et al.
Veröffentlicht: (2025)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
Synthetic Speech Classification: IEEE Signal Processing Cup 2022 challenge
von: Rahmun, Mahieyin, et al.
Veröffentlicht: (2024)
von: Rahmun, Mahieyin, et al.
Veröffentlicht: (2024)
IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks
von: Lee, Sunghwa, et al.
Veröffentlicht: (2025)
von: Lee, Sunghwa, et al.
Veröffentlicht: (2025)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
von: Li, Xingchen, et al.
Veröffentlicht: (2025)
von: Li, Xingchen, et al.
Veröffentlicht: (2025)
SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
von: Im, Jaekwon, et al.
Veröffentlicht: (2025)
von: Im, Jaekwon, et al.
Veröffentlicht: (2025)
Modeling Multi-Level Hearing Loss for Speech Intelligibility Prediction
von: Zhou, Xiajie, et al.
Veröffentlicht: (2025)
von: Zhou, Xiajie, et al.
Veröffentlicht: (2025)
Speech Separation using Neural Audio Codecs with Embedding Loss
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution
von: Abreu, Wallace, et al.
Veröffentlicht: (2026)
von: Abreu, Wallace, et al.
Veröffentlicht: (2026)
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026)
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026)
Combined Generative and Predictive Modeling for Speech Super-resolution
von: Wang, Heming, et al.
Veröffentlicht: (2024)
von: Wang, Heming, et al.
Veröffentlicht: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2024)
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2024)
A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
von: Nan, Zheng, et al.
Veröffentlicht: (2024)
von: Nan, Zheng, et al.
Veröffentlicht: (2024)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024)
von: Lee, Joun Yeop, et al.
Veröffentlicht: (2024)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
von: Bijoy, Mehedi Hasan, et al.
Veröffentlicht: (2025)
von: Bijoy, Mehedi Hasan, et al.
Veröffentlicht: (2025)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss
von: Liang, Yayun, et al.
Veröffentlicht: (2026)
von: Liang, Yayun, et al.
Veröffentlicht: (2026)
AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement
von: Opochinsky, Renana, et al.
Veröffentlicht: (2026)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2026)
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WaLi: Can Pressure Sensors in HVAC Systems Capture Human Speech?
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025) -
CIS-BWE: Chaos-Informed Speech Bandwidth Extension
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025) -
HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025) -
SUBARU: A Practical Approach to Power Saving in Hearables Using SUB-Nyquist Audio Resolution Upsampling
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025) -
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
von: Hasan, Rashedul, et al.
Veröffentlicht: (2025)