CIS-BWE: Chaos-Informed Speech Bandwidth Extension
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tamiti, Tarikul Islam, Das, Tonmoy, Mamun, Nursadul, Barua, Anomadarshi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
WaLi: Can Pressure Sensors in HVAC Systems Capture Human Speech?
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
NLDSI-BWE: Non Linear Dynamical Systems-Inspired Multi Resolution Discriminators for Speech Bandwidth Extension
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
SUBARU: A Practical Approach to Power Saving in Hearables Using SUB-Nyquist Audio Resolution Upsampling
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
par: Tamiti, Tarikul Islam, et autres
Publié: (2025)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
par: Hasan, Rashedul, et autres
Publié: (2025)
par: Hasan, Rashedul, et autres
Publié: (2025)
EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network
par: Avro, Shamin Bin Habib, et autres
Publié: (2025)
par: Avro, Shamin Bin Habib, et autres
Publié: (2025)
Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks
par: Salhab, Mahmoud, et autres
Publié: (2024)
par: Salhab, Mahmoud, et autres
Publié: (2024)
UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension
par: Gupta, Kishan, et autres
Publié: (2025)
par: Gupta, Kishan, et autres
Publié: (2025)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
par: Fang, Yuan, et autres
Publié: (2024)
par: Fang, Yuan, et autres
Publié: (2024)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
par: Lu, Ye-Xin, et autres
Publié: (2024)
par: Lu, Ye-Xin, et autres
Publié: (2024)
A Fly on the Wall -- Exploiting Acoustic Side-Channels in Differential Pressure Sensors
par: Achamyeleh, Yonatan Gizachew, et autres
Publié: (2024)
par: Achamyeleh, Yonatan Gizachew, et autres
Publié: (2024)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
par: Das, Shoutrik, et autres
Publié: (2025)
par: Das, Shoutrik, et autres
Publié: (2025)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
par: Zhang, Xu, et autres
Publié: (2026)
par: Zhang, Xu, et autres
Publié: (2026)
Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses
par: Ghosh, Suhita, et autres
Publié: (2024)
par: Ghosh, Suhita, et autres
Publié: (2024)
Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing
par: Liu, Tianchi, et autres
Publié: (2025)
par: Liu, Tianchi, et autres
Publié: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
par: Lin, Zijian, et autres
Publié: (2025)
par: Lin, Zijian, et autres
Publié: (2025)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
par: Kim, Jaeyoung, et autres
Publié: (2024)
par: Kim, Jaeyoung, et autres
Publié: (2024)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
par: Wang, Yongqi, et autres
Publié: (2023)
par: Wang, Yongqi, et autres
Publié: (2023)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
par: Bhuiyan, Mohammed Aman, et autres
Publié: (2026)
par: Bhuiyan, Mohammed Aman, et autres
Publié: (2026)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
par: Wang, Helin, et autres
Publié: (2025)
par: Wang, Helin, et autres
Publié: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
par: Lee, Seo-Hyun, et autres
Publié: (2023)
par: Lee, Seo-Hyun, et autres
Publié: (2023)
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
par: Ji, Zhoulin, et autres
Publié: (2024)
par: Ji, Zhoulin, et autres
Publié: (2024)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
par: Lau, Hok-Shing, et autres
Publié: (2024)
par: Lau, Hok-Shing, et autres
Publié: (2024)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
par: Shi, Jiatong, et autres
Publié: (2025)
par: Shi, Jiatong, et autres
Publié: (2025)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
par: Bian, Weizhen, et autres
Publié: (2024)
par: Bian, Weizhen, et autres
Publié: (2024)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
par: Moell, Birger, et autres
Publié: (2025)
par: Moell, Birger, et autres
Publié: (2025)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
par: Shah, Neil, et autres
Publié: (2024)
par: Shah, Neil, et autres
Publié: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
par: Choi, Yerin, et autres
Publié: (2024)
par: Choi, Yerin, et autres
Publié: (2024)
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
par: Songyi, Li, et autres
Publié: (2026)
par: Songyi, Li, et autres
Publié: (2026)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
par: Lu, Ye-Xin, et autres
Publié: (2024)
par: Lu, Ye-Xin, et autres
Publié: (2024)
TinyML for Speech Recognition
par: Barovic, Andrew, et autres
Publié: (2025)
par: Barovic, Andrew, et autres
Publié: (2025)
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
par: Shah, Neil, et autres
Publié: (2024)
par: Shah, Neil, et autres
Publié: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
par: Shi, Hao, et autres
Publié: (2024)
par: Shi, Hao, et autres
Publié: (2024)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
par: Wang, Helin, et autres
Publié: (2025)
par: Wang, Helin, et autres
Publié: (2025)
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
par: Zhou, Haoshuai, et autres
Publié: (2025)
par: Zhou, Haoshuai, et autres
Publié: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
par: Wang, Xinsheng, et autres
Publié: (2025)
par: Wang, Xinsheng, et autres
Publié: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
par: Anastassiou, Philip, et autres
Publié: (2024)
par: Anastassiou, Philip, et autres
Publié: (2024)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
par: Yao, Wenhan, et autres
Publié: (2024)
par: Yao, Wenhan, et autres
Publié: (2024)
Hello Afrika: Speech Commands in Kinyarwanda
par: Igwegbe, George, et autres
Publié: (2025)
par: Igwegbe, George, et autres
Publié: (2025)
DASB - Discrete Audio and Speech Benchmark
par: Mousavi, Pooneh, et autres
Publié: (2024)
par: Mousavi, Pooneh, et autres
Publié: (2024)
Documents similaires
-
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
par: Tamiti, Tarikul Islam, et autres
Publié: (2025) -
WaLi: Can Pressure Sensors in HVAC Systems Capture Human Speech?
par: Tamiti, Tarikul Islam, et autres
Publié: (2025) -
NLDSI-BWE: Non Linear Dynamical Systems-Inspired Multi Resolution Discriminators for Speech Bandwidth Extension
par: Tamiti, Tarikul Islam, et autres
Publié: (2025) -
SUBARU: A Practical Approach to Power Saving in Hearables Using SUB-Nyquist Audio Resolution Upsampling
par: Tamiti, Tarikul Islam, et autres
Publié: (2025) -
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
par: Hasan, Rashedul, et autres
Publié: (2025)